Yotta
Research Engineer Intern - AI Systems
- Confirmed live in the last 24 hours
- No salary listed
- Internship
- Remote · United States
- Added 2 months ago
About this role
Yotta Labs is seeking a Research Engineer Intern to contribute to optimizing AI systems and infrastructure, particularly focusing on Trainium, GPU kernels, and LLM performance. This 12-16 week internship offers a chance to own a project from design to implementation, with potential for a full-time return offer based on performance. You'll work remotely and collaboratively, impacting the performance of AI applications on a cutting-edge platform.
What you'll do
- Implement and optimize compute kernels for AI operations (Attention, GEMM, MoE, quantization).
- Develop custom operators using CUDA, Triton, ROCm/HIP, or Neuron SDK with PyTorch/XLA.
- Profile and optimize inference performance within vLLM, SGLang, and custom runtimes.
- Build benchmarks and debug performance issues.
- Contribute code to open-source AI infrastructure projects.
- Ship code with tests and documentation.
Benefits
- Remote work environment
- Access to cutting-edge hardware
- Competitive internship compensation
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Yotta
Likely interview questions
- Describe a time you optimized code for performance. What tools did you use, and what were the key improvements?
- Explain your understanding of GPU memory hierarchy and how it impacts performance optimization.