Skip to main content

OpenAI

Software Engineer, RL Training Infra

San Francisco (Remote)$295k–$445kfulltimemidAdded 1 month ago

About this role

Join OpenAI's Post-Training Frontiers team to keep large-scale reinforcement learning training runs operational and efficient. You'll solve critical engineering and infrastructure challenges across training systems, inference, and distributed infrastructure while supporting the development of frontier AI agents shipped in products like ChatGPT and the API.

What you'll do

  • Debug and resolve urgent engineering and infrastructure issues blocking RL training runs
  • Troubleshoot failures across training systems, inference, orchestration, scaling, and distributed infrastructure
  • Improve reliability and efficiency of large-scale model training pipelines
  • Support researchers developing infrastructure-heavy capabilities like multi-agent systems and memory
  • Convert recurring operational problems into robust tools, systems, and processes
  • Collaborate with research and infrastructure teams during tight model training timelines

What they're looking for

  • ML infrastructure (training systems, RL, or inference)
  • Distributed systems debugging across GPUs and networking
  • Scaling and orchestration systems
  • Performance optimization and production infrastructure
  • Fast learning and cross-layer problem-solving
  • Deep debugging and root cause analysis
  • Strong ownership and communication
  • Experience with large-scale model training (nice to have)
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

OpenAI

OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.

View all jobs at OpenAI

Likely interview questions

  • Walk us through a time you debugged a complex issue that spanned multiple layers of a system (e.g., training, inference, or infrastructure). How did you approach it?
  • Describe your experience with reinforcement learning systems or training infrastructure. What specific RL or ML infrastructure problems have you solved?