Lumaai
Research Scientist / Engineer – Reinforcement Learning Infrastructure
- Confirmed live in the last 24 hours
- $195k–$395k
- Mid level
- Full-time
- Remote · Redwood City, CA
- Added 2 months ago
About this role
Luma is seeking a Research Scientist/Engineer to build and scale reinforcement learning infrastructure, enabling their models to become truly useful. This role focuses on designing and implementing distributed systems that handle training, rollout generation, environment execution, and reward computation at a massive scale, utilizing thousands of GPUs. The ideal candidate will have extensive experience post-training LLMs with RL and debugging asynchronous rollout pipelines.
What you'll do
- Design, build, and scale distributed RL post-training systems.
- Build high-throughput rollout generation pipelines.
- Design and build scalable RL environments for agentic tasks.
- Develop reward infrastructure including verifiable rewards and LLM-as-judge pipelines.
- Develop evaluation, monitoring, and debugging tooling.
- Collaborate with researchers to advance training efficiency and stability.
What they're looking for
- Reinforcement Learning (RL)
- Distributed PyTorch
- LLMs
- PPO/GRPO
- RLHF
- vLLM
- SGLang
- NCCL
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Lumaai
- Industry
- Technology & Software
Likely interview questions
- Describe your experience post-training LLMs with RL at scale.
- How familiar are you with FSDP, Tensor/Pipeline/Expert Parallel?