Skip to main content

OpenAI

Software Engineer, Infrastructure

San Francisco$210k–$405kfulltimemidAdded today

About this role

Join OpenAI's Infrastructure team to design and operate mission-critical distributed systems supporting ChatGPT and the OpenAI API. You'll collaborate across teams to build scalable, reliable infrastructure spanning core systems, observability, reliability engineering, and cloud infrastructure.

What you'll do

  • Design, build, and maintain highly available and performant distributed systems used across the engineering organization
  • Define technical strategy, architecture, and long-term goals in collaboration with your team
  • Partner with engineers, product managers, and researchers to evolve infrastructure for emerging needs
  • Enhance internal tooling, automation, and developer experience
  • Lead incident response, postmortems, and establish best practices for system reliability and scalability

What they're looking for

  • Proficiency in Python, Go, C++, Rust, or similar languages
  • Experience designing or operating distributed systems at scale
  • Linux administration and containerization (Kubernetes)
  • Infrastructure-as-code tools (Terraform)
  • CI/CD pipeline design and implementation
  • Observability stacks (metrics, logging, tracing)
  • Systems debugging and performance optimization
  • Cross-functional communication and technical leadership
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

OpenAI

OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.

View all jobs at OpenAI

Likely interview questions

  • Describe a complex distributed system you've designed or scaled—what were the key reliability and performance challenges?
  • How do you approach debugging performance bottlenecks in large-scale systems, and what tools have you found most effective?