OpenAI
Software Engineer, Compute Foundations
About this role
OpenAI seeks a Software Engineer to design and operate Kubernetes-based distributed systems that manage GPU compute infrastructure across data centers and providers. You'll build controllers, APIs, and provisioning services that handle the complete lifecycle of bare-metal machines—from boot through configuration, upgrades, and decommissioning—while ensuring reliability as the fleet scales.
What you'll do
- Design and build Kubernetes controllers and distributed services coordinating infrastructure across multiple sites with failure isolation and scaling
- Define APIs and resource models for lifecycle operations across heterogeneous hardware platforms and providers
- Develop provisioning and configuration services integrating network boot, BMCs, firmware, OS images, drivers, and host configuration
- Build lifecycle management for discovery, allocation, provisioning, upgrades, maintenance, recovery, and decommissioning
- Design reliable reconciliation and recovery handling concurrent changes, interrupted operations, and partial failures
- Optimize control-plane throughput, API latency, and convergence time while respecting external system and provider rate limits
What they're looking for
- Distributed systems design and production operations
- Kubernetes controllers and reconciliation patterns
- Bare-metal provisioning (PXE, DHCP, DNS, BMCs, firmware, Linux)
- API design with focus on concurrency, consistency, idempotency, and failure handling
- Systems diagnosis across service, OS, and hardware boundaries
- Infrastructure control planes and multi-site coordination
- GPU or HPC infrastructure experience
- Cross-team communication on technical tradeoffs
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
OpenAI
OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.
View all jobs at OpenAILikely interview questions
- Describe a time you designed a distributed system API that had to remain consistent across multiple failure domains—what tradeoffs did you make?
- Walk us through how you'd diagnose a provisioning failure that could originate in network boot, BMC communication, or image deployment.