OpenAI
Software Engineer, RL Training Infra
About this role
Join OpenAI's Post-Training Frontiers team to keep large-scale reinforcement learning training runs operational and efficient. You'll solve critical engineering and infrastructure challenges across training systems, inference, and distributed infrastructure while supporting the development of frontier AI agents shipped in products like ChatGPT and the API.
What you'll do
- Debug and resolve urgent engineering and infrastructure issues blocking RL training runs
- Troubleshoot failures across training systems, inference, orchestration, scaling, and distributed infrastructure
- Improve reliability and efficiency of large-scale model training pipelines
- Support researchers developing infrastructure-heavy capabilities like multi-agent systems and memory
- Convert recurring operational problems into robust tools, systems, and processes
- Collaborate with research and infrastructure teams during tight model training timelines
What they're looking for
- ML infrastructure (training systems, RL, or inference)
- Distributed systems debugging across GPUs and networking
- Scaling and orchestration systems
- Performance optimization and production infrastructure
- Fast learning and cross-layer problem-solving
- Deep debugging and root cause analysis
- Strong ownership and communication
- Experience with large-scale model training (nice to have)
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
OpenAI
OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.
View all jobs at OpenAILikely interview questions
- Walk us through a time you debugged a complex issue that spanned multiple layers of a system (e.g., training, inference, or infrastructure). How did you approach it?
- Describe your experience with reinforcement learning systems or training infrastructure. What specific RL or ML infrastructure problems have you solved?