OpenAI
Machine Learning Engineer, Distributed Data Systems - Robotics
About this role
OpenAI seeks a Machine Learning Engineer to design and scale distributed data infrastructure for large-scale multimodal training and evaluation in robotics. You'll build robust pipelines, collaborate with researchers, and maintain critical systems supporting OpenAI's rapid iteration cycles.
What you'll do
- Design and build distributed compute, data orchestration, and storage systems
- Scale data platforms reliably while maintaining efficiency and security
- Partner with researchers to translate requirements into production systems
- Harden, optimize, and maintain critical multimodal training infrastructure
- Ensure data pipelines meet scalability and reliability requirements
- Support evaluation systems for robotics and AI research
What they're looking for
- Distributed systems architecture
- Large-scale infrastructure design
- Data pipeline orchestration
- Software engineering fundamentals
- System reliability and hardening
- Distributed storage technologies
- Problem-solving under ambiguity
- Cross-functional collaboration
Benefits
- Hybrid work model (3 days in office per week)
- Based in San Francisco, CA
- Relocation assistance provided
- Work on cutting-edge robotics and AGI research
- Collaborative environment with leading researchers
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
OpenAI
OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.
View all jobs at OpenAILikely interview questions
- Describe your experience designing and maintaining distributed data infrastructure at scale. What were the biggest challenges you faced, and how did you approach reliability and efficiency?
- Tell us about a time you had to translate complex research requirements into a production-ready system. How did you handle ambiguity and changing requirements?