Together AI
Systems Research Engineer Intern - GPU Programming (Summer 2027)
About this role
Join Together AI as a GPU Programming Systems Research Engineer Intern this summer, optimizing GPU-accelerated kernels and algorithms for AI/ML applications. You'll collaborate with modeling, algorithm, hardware, and software teams to co-design efficient GPU architectures while staying current with cutting-edge GPU programming techniques.
What you'll do
- Develop and optimize GPU-accelerated kernels for ML/AI applications
- Fine-tune GPU code for improved performance and scalability
- Co-design GPU kernels with modeling and algorithm teams
- Integrate GPU-accelerated solutions into existing software systems
- Research and evaluate latest GPU programming advancements
- Profile and optimize GPU code using performance tools
What they're looking for
- GPU programming (CUDA and/or Triton)
- Parallel computing
- ML/AI model knowledge
- Performance profiling and optimization tools
- Problem-solving and analytical thinking
- Cross-functional collaboration
- Systems architecture understanding
- C++ or Python
Benefits
- Competitive hourly compensation ($58–$63/hour)
- Housing stipend
- 12–14 week structured internship program
- Work with industry-leading engineers and researchers
- Access to cutting-edge GPU and AI infrastructure
- Flexible internship dates (May or June start)
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Together AI
Together AI builds GPU compute infrastructure and open-source model customization platforms for AI developers and enterprises. The company is hiring for infrastructure operations, ML systems engineering, go-to-market technology, customer success, and GPU research roles.
- Website
- together.ai
Likely interview questions
- Describe your experience optimizing CUDA or Triton kernels. What performance metrics did you improve and how?
- Walk us through how you'd approach profiling and identifying bottlenecks in a GPU-accelerated ML model.