OpenAI
Backend Software Engineer (Evals)
About this role
Join OpenAI's Support Automation team to design and build evaluation infrastructure for AI-powered support systems. You'll develop backend services and evaluation pipelines that measure quality and enable continuous monitoring of support automation at scale, working closely with data science and research teams.
What you'll do
- Design reliable, reproducible, and extensible evaluation pipelines for support automation
- Build infrastructure for continuous monitoring frameworks including regression detection and golden dataset management
- Develop backend services and APIs supporting intelligent automation and knowledge systems
- Integrate and structure data across platforms for downstream AI workflows
- Collaborate with data, research, and engineering teams to integrate OpenAI models into high-leverage workflows
- Own the full development lifecycle of new backend systems and internal platform capabilities
What they're looking for
- Backend engineering (4+ years at product companies)
- Python, FastAPI, and PostgreSQL
- Distributed systems and API design
- ML/LLM evaluation methods and metrics
- AI agent and application development
- Data pipeline design and scaling
- Production ML systems and performance measurement
- Pragmatic iterative development approach
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
OpenAI
OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.
View all jobs at OpenAILikely interview questions
- Walk us through your experience building evaluation frameworks for ML/LLM systems. What metrics did you track and how did you ensure they were reliable and reproducible at scale?
- Describe a time you designed a distributed system or data pipeline. What were the key architectural decisions you made and how did you handle scalability?