OpenAI
Software Engineer, Infrastructure
About this role
Join OpenAI's Infrastructure team to design and operate mission-critical distributed systems supporting ChatGPT and the OpenAI API. You'll collaborate across teams to build scalable, reliable infrastructure spanning core systems, observability, reliability engineering, and cloud infrastructure.
What you'll do
- Design, build, and maintain highly available and performant distributed systems used across the engineering organization
- Define technical strategy, architecture, and long-term goals in collaboration with your team
- Partner with engineers, product managers, and researchers to evolve infrastructure for emerging needs
- Enhance internal tooling, automation, and developer experience
- Lead incident response, postmortems, and establish best practices for system reliability and scalability
What they're looking for
- Proficiency in Python, Go, C++, Rust, or similar languages
- Experience designing or operating distributed systems at scale
- Linux administration and containerization (Kubernetes)
- Infrastructure-as-code tools (Terraform)
- CI/CD pipeline design and implementation
- Observability stacks (metrics, logging, tracing)
- Systems debugging and performance optimization
- Cross-functional communication and technical leadership
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
OpenAI
OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.
View all jobs at OpenAILikely interview questions
- Describe a complex distributed system you've designed or scaled—what were the key reliability and performance challenges?
- How do you approach debugging performance bottlenecks in large-scale systems, and what tools have you found most effective?