Skip to main content

OpenAI

Software Engineer, Platform Systems

San Francisco$310k–$460kfulltimemidAdded 1 month ago

About this role

Join OpenAI's Platform Systems team to build distributed systems that monitor and operate large-scale AI training workloads. You'll develop failure detection, tracing, and observability infrastructure that enables reliable training on massive supercomputers, working at the core of OpenAI's training platform.

What you'll do

  • Design and build distributed failure detection, tracing, and profiling systems for large-scale AI training
  • Develop tooling to identify and diagnose slow, faulty, or misbehaving nodes in distributed systems
  • Improve observability, reliability, and performance across the training platform
  • Debug and resolve issues in complex, high-throughput distributed systems
  • Collaborate with systems, infrastructure, and research teams to evolve platform capabilities
  • Adapt failure detection and tracing systems to support new training paradigms and workloads

What they're looking for

  • Distributed systems design and debugging
  • Low-level systems programming and performance optimization
  • Hardware, operating systems, and networking knowledge
  • Concurrency and multi-threaded programming
  • High-performance computing experience
  • Observability and monitoring systems
  • Failure detection and tracing systems
  • System-level troubleshooting and automation
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

OpenAI

OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.

View all jobs at OpenAI

Likely interview questions

  • Describe your experience designing or debugging distributed systems at scale. What was the largest system you've worked on, and what made it challenging?
  • How would you approach building a failure detection system for a cluster of thousands of nodes? What metrics or signals would you prioritize?