OpenAI
Simulation Infrastructure Engineer
About this role
OpenAI is seeking a Simulation Applications Engineer to build production-grade pipelines that automate robotics model training, evaluation, and hardware validation. You'll own CI/CD systems, orchestration infrastructure, and metrics tooling that enable large-scale simulation runs and seamless hardware-in-the-loop integration.
What you'll do
- Build and maintain CI/CD pipelines and presubmit checks for simulation code, environments, and tasks
- Implement end-to-end automation for model evaluation in simulation (SIL) and hardware-in-the-loop (HIL) runs with metric computation and reporting
- Design APIs and connectors for research and training systems to schedule, seed, and batch simulations
- Build scheduling and orchestration infrastructure to run tens of thousands of concurrent rollouts with GPU optimization
- Create metrics, dashboards, and presubmit tests to detect simulation health regressions and measure fidelity
- Implement artifact versioning, environment immutability, experiment provenance, and resource quota management
What they're looking for
- CI/CD pipelines and infrastructure at scale
- Distributed compute and cloud GPU workload orchestration
- Hardware-in-the-loop (HIL) and software-in-the-loop (SIL) systems
- Python, C++, or Rust
- Container orchestration (Kubernetes) and distributed task queues
- System observability, metrics, and alerting
- API design and developer ergonomics
- Reinforcement learning tooling or large-scale data pipelines (bonus)
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
OpenAI
OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.
View all jobs at OpenAILikely interview questions
- Walk us through a CI/CD pipeline you've built at scale—what were the key challenges in making it reliable for thousands of concurrent jobs, and how did you handle failures?
- Describe your experience orchestrating distributed compute workloads (GPUs, containers, task queues). How did you optimize for throughput and cost in a large-scale system?