Together AI
Systems Research Engineer Intern - GPU Programming (Winter 2027)
San FranciscointernshipinternAdded today
About this role
Develop and optimize GPU-accelerated kernels for ML/AI applications at Together AI's San Francisco headquarters. You'll collaborate across teams to co-design efficient GPU architectures and programming models while staying current with the latest GPU optimization techniques.
What you'll do
- Optimize and fine-tune GPU code for improved performance and scalability
- Co-design GPU kernels and model architectures with the modeling and algorithm team
- Integrate GPU-accelerated solutions into existing software systems with cross-functional teams
- Profile and benchmark GPU performance using industry-standard tools
- Research and implement latest GPU programming techniques and technologies
- Contribute to efficient GPU architecture and programming model design
What they're looking for
- GPU programming (CUDA and/or Triton)
- Parallel computing and optimization
- ML/AI applications and model architecture knowledge
- Performance profiling and optimization tools
- Problem-solving and analytical thinking
- C++ or Python for systems programming
- Hardware-software co-design principles
- Large-scale inference and training systems
Benefits
- Competitive hourly compensation ($58–$63/hour)
- Housing stipend
- On-site work at San Francisco HQ
- 12–14 week structured internship program
- Mentorship from industry-leading engineers and researchers
- Exposure to cutting-edge AI infrastructure at scale
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Together AI
Together AI builds GPU compute infrastructure and open-source model customization platforms for AI developers and enterprises. The company is hiring for infrastructure operations, ML systems engineering, go-to-market technology, customer success, and GPU research roles.
- Website
- together.ai
Likely interview questions
- Walk us through a GPU kernel optimization you've done—what bottleneck did you identify and how did you address it?
- How do you approach profiling GPU code to identify performance issues?