Skip to main content

OpenAI

Backend Software Engineer (Evals)

San Francisco$230k–$385kfulltimemidAdded 1 month ago

About this role

Join OpenAI's Support Automation team to design and build evaluation infrastructure for AI-powered support systems. You'll develop backend services and evaluation pipelines that measure quality and enable continuous monitoring of support automation at scale, working closely with data science and research teams.

What you'll do

  • Design reliable, reproducible, and extensible evaluation pipelines for support automation
  • Build infrastructure for continuous monitoring frameworks including regression detection and golden dataset management
  • Develop backend services and APIs supporting intelligent automation and knowledge systems
  • Integrate and structure data across platforms for downstream AI workflows
  • Collaborate with data, research, and engineering teams to integrate OpenAI models into high-leverage workflows
  • Own the full development lifecycle of new backend systems and internal platform capabilities

What they're looking for

  • Backend engineering (4+ years at product companies)
  • Python, FastAPI, and PostgreSQL
  • Distributed systems and API design
  • ML/LLM evaluation methods and metrics
  • AI agent and application development
  • Data pipeline design and scaling
  • Production ML systems and performance measurement
  • Pragmatic iterative development approach
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

OpenAI

OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.

View all jobs at OpenAI

Likely interview questions

  • Walk us through your experience building evaluation frameworks for ML/LLM systems. What metrics did you track and how did you ensure they were reliable and reproducible at scale?
  • Describe a time you designed a distributed system or data pipeline. What were the key architectural decisions you made and how did you handle scalability?