OpenAI
Training - Runtime Foundations Engineer
About this role
Build and optimize the distributed runtime system that orchestrates massive-scale machine learning training across thousands of machines. You'll develop high-performance, fault-tolerant software in Rust and Python that enables researchers to run experiments and frontier-scale model training reliably while maintaining system stability and performance.
What you'll do
- Design and build software to orchestrate ML workloads across large supercomputer clusters
- Profile and optimize the stack to support computation at frontier scale
- Improve reliability, observability, and fault tolerance for long-running distributed jobs
- Debug complex distributed systems issues across large clusters
- Work across Python and Rust codebase for runtime foundations
- Respond to evolving ML system needs to support researcher requirements
What they're looking for
- Distributed systems design and development
- Rust programming language
- Asynchronous and concurrent system programming
- Linux systems administration and debugging
- Performance analysis and memory profiling
- High-performance computing optimization
- Python
- Systems-level debugging and troubleshooting
Benefits
- Hybrid work model (3 days in office per week)
- Based in San Francisco with relocation assistance available
- Work on frontier-scale AI infrastructure
- High ownership environment with light process
- Strong engineering agency and autonomy
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
OpenAI
OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.
View all jobs at OpenAILikely interview questions
- Tell us about a distributed system you've debugged—what was the issue and how did you approach finding the root cause?
- Describe your experience optimizing system performance at scale. What tools and techniques do you rely on?