Skip to main content

OpenAI

Training - Runtime Foundations Engineer

San Francisco (Remote)$266k–$385kfulltimemidAdded 2 days ago

About this role

Build and optimize the distributed runtime system that orchestrates massive-scale machine learning training across thousands of machines. You'll develop high-performance, fault-tolerant software in Rust and Python that enables researchers to run experiments and frontier-scale model training reliably while maintaining system stability and performance.

What you'll do

  • Design and build software to orchestrate ML workloads across large supercomputer clusters
  • Profile and optimize the stack to support computation at frontier scale
  • Improve reliability, observability, and fault tolerance for long-running distributed jobs
  • Debug complex distributed systems issues across large clusters
  • Work across Python and Rust codebase for runtime foundations
  • Respond to evolving ML system needs to support researcher requirements

What they're looking for

  • Distributed systems design and development
  • Rust programming language
  • Asynchronous and concurrent system programming
  • Linux systems administration and debugging
  • Performance analysis and memory profiling
  • High-performance computing optimization
  • Python
  • Systems-level debugging and troubleshooting

Benefits

  • Hybrid work model (3 days in office per week)
  • Based in San Francisco with relocation assistance available
  • Work on frontier-scale AI infrastructure
  • High ownership environment with light process
  • Strong engineering agency and autonomy
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

OpenAI

OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.

View all jobs at OpenAI

Likely interview questions

  • Tell us about a distributed system you've debugged—what was the issue and how did you approach finding the root cause?
  • Describe your experience optimizing system performance at scale. What tools and techniques do you rely on?