Skip to main content

Prime Intellect

Research Engineer - RL Infrastructure

San FranciscofulltimemidAdded 1 month ago

About this role

Prime Intellect seeks a Research Engineer to enhance their large-scale reinforcement learning (RL) infrastructure, focusing on optimizing performance, memory efficiency, and system throughput. The role involves collaborating with engineers and researchers to improve training systems and contribute to the architectural design for RL training.

What you'll do

  • Build and enhance systems for large-scale RL training
  • Optimize training efficiency across various layers
  • Implement performance optimizations
  • Work on distributed training systems
  • Contribute to open-source libraries
  • Collaborate on systems improvements

What they're looking for

  • Systems engineering in AI/ML
  • Experience with PyTorch and distributed frameworks
  • Performance optimization skills
  • Knowledge of large-scale training techniques
  • Understanding of GPU architecture
  • Ability to identify and resolve bottlenecks
  • Comfort in fast-paced environments
  • Experience with CUDA / Triton kernels (plus)

Benefits

  • Competitive cash compensation and equity
  • Flexible remote work options
  • Visa sponsorship and relocation support
  • Quarterly team events and learning opportunities
  • Work with a highly technical team
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Prime Intellect

Prime Intellect builds AI infrastructure and training systems, with a focus on reinforcement learning, distributed training, and compute optimization. The company is hiring Research Engineers and Infrastructure Engineers to develop synthetic data pipelines, decentralized training infrastructure, compute visibility systems, and large-scale RL training optimization.

View all jobs at Prime Intellect

Likely interview questions

  • Walk us through a time you optimized a large-scale training workload. What was the bottleneck, how did you identify it, and what was the impact?
  • Describe your experience with distributed training frameworks like PyTorch Distributed, DeepSpeed, or Megatron. Which have you worked with and what trade-offs did you encounter?