Skip to main content

OpenAI

Applied AI Engineer, Codex Core Agent

San Francisco$230k–$385kfulltimemidAdded 1 month ago

About this role

Join OpenAI's Codex Core Agent team to transform AI research into production-ready agent systems that solve real software engineering tasks. You'll bridge research and product by improving agent performance, reliability, and user value through evaluation, experimentation, and systematic failure analysis.

What you'll do

  • Design and iterate on agent behaviors for real-world coding tasks and long-horizon workflows
  • Develop evaluations to measure agent performance, identify failures, and detect regressions
  • Optimize agent performance through prompting, tool-use strategies, and context construction
  • Analyze production failures and improve robustness and reliability systematically
  • Build feedback loops and data systems for better real-task data in evaluation
  • Collaborate with product teams on user-facing agent experiences and interfaces

What they're looking for

  • Python programming
  • Machine learning and LLM product development
  • Model evaluation and prompt design
  • Agent frameworks and tool-using LLM systems
  • Systems thinking and debugging complex failures
  • Data analysis and handling large datasets
  • Fine-tuning and model experimentation
  • Code generation or developer tooling experience
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

OpenAI

OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.

View all jobs at OpenAI

Likely interview questions

  • Walk us through a time you shipped an LLM or ML-powered feature. What metrics did you use to measure success, and what gap did you find between lab performance and real-world usefulness?
  • Tell us about a time you debugged agent or model failure in production. How did you systematically identify the root cause and turn it into a product improvement?