Skip to main content

Yotta

Research Engineer Intern - AI Systems

  • Confirmed live in the last 24 hours
  • No salary listed
  • Internship
  • Remote · United States
  • Added 2 months ago

About this role

Yotta Labs is seeking a Research Engineer Intern to contribute to optimizing AI systems and infrastructure, particularly focusing on Trainium, GPU kernels, and LLM performance. This 12-16 week internship offers a chance to own a project from design to implementation, with potential for a full-time return offer based on performance. You'll work remotely and collaboratively, impacting the performance of AI applications on a cutting-edge platform.

What you'll do

  • Implement and optimize compute kernels for AI operations (Attention, GEMM, MoE, quantization).
  • Develop custom operators using CUDA, Triton, ROCm/HIP, or Neuron SDK with PyTorch/XLA.
  • Profile and optimize inference performance within vLLM, SGLang, and custom runtimes.
  • Build benchmarks and debug performance issues.
  • Contribute code to open-source AI infrastructure projects.
  • Ship code with tests and documentation.

Benefits

  • Remote work environment
  • Access to cutting-edge hardware
  • Competitive internship compensation
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Yotta

View all jobs at Yotta

Likely interview questions

  • Describe a time you optimized code for performance. What tools did you use, and what were the key improvements?
  • Explain your understanding of GPU memory hierarchy and how it impacts performance optimization.