OpenAI
AI Infrastructure Engineer, pAGI
About this role
Join OpenAI's pAGI Infra team to design and operate large-scale systems for AI model training and evaluation. You'll tackle distributed systems challenges, GPU optimization, and infrastructure tooling while partnering with researchers to accelerate the path from experiment to production.
What you'll do
- Build and operate infrastructure for large-scale training and evaluation workflows, optimizing reliability and resource efficiency
- Develop shared inference and grading platforms with automated capacity management and performance monitoring
- Improve compute scheduling and resource allocation to reduce GPU idle time and improve failure recovery
- Diagnose performance bottlenecks across training, inference, and orchestration layers
- Create self-service tools and observability solutions to help researchers manage experiments with minimal manual intervention
- Partner with research and engineering teams to translate requirements into dependable infrastructure
What they're looking for
- Distributed systems design and operation
- GPU performance optimization
- ML infrastructure and inference systems
- Compute scheduling and resource allocation
- Performance debugging and measurement
- Infrastructure tooling and automation
- Software engineering fundamentals
- Observability and monitoring systems
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
OpenAI
OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.
View all jobs at OpenAILikely interview questions
- Describe your experience building or operating distributed systems at scale—what were the key challenges and how did you address them?
- Tell us about a time you identified and resolved a critical performance bottleneck in a production system. What tools and metrics did you use?