Skip to main content

OpenAI

Software Engineer, ML Systems & Training Architecture

San Francisco$295k–$380kfulltimemidAdded 1 month ago

About this role

OpenAI seeks a Senior Software Engineer to strengthen their robotics team's ML training infrastructure. You'll maintain and improve training frameworks, debug complex ML systems, and unblock researchers by ensuring code quality and infrastructure reliability while working 5 days/week in San Francisco.

What you'll do

  • Review, improve, and refactor code across training frameworks and infrastructure
  • Identify and prevent risky or low-quality changes through code review
  • Debug issues spanning ML training systems, GPUs, clusters, and networking
  • Unblock researchers and engineers from broken training jobs and workflow failures
  • Enhance reliability, maintainability, and usability of the training framework
  • Ship practical engineering solutions that directly improve team velocity

What they're looking for

  • Strong software engineering fundamentals and code review expertise
  • ML systems, training frameworks, or distributed systems experience
  • GPU and infrastructure debugging
  • Ability to read and debug unfamiliar codebases quickly
  • High-velocity shipping with pragmatic judgment
  • Experience with messy or fast-moving codebases
  • Root cause analysis and problem-solving

Benefits

  • Compensation: $295K-$380K USD
  • Relocation assistance provided
  • Work on cutting-edge robotics and AGI research
  • 5 days/week in-office (San Francisco)
  • Collaborative environment with researchers and engineers
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

OpenAI

OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.

View all jobs at OpenAI

Likely interview questions

  • Walk us through a time you debugged a complex issue across ML training systems, GPUs, or distributed infrastructure. What was your approach and how did you get to root cause?
  • Describe your experience with ML training frameworks and infrastructure. What specific frameworks or systems have you worked with, and what improvements did you make to them?