OpenAI
Software Engineer, API Frontiers
About this role
Build and operate backend services powering OpenAI's Responses API, connecting frontier AI models to developers. You'll design production APIs, manage staged rollouts, and ensure reliability for long-running agent workflows while collaborating with research and safety teams.
What you'll do
- Design, build, and operate backend services that expose frontier model capabilities through developer-facing APIs
- Partner with Research, Safety, and API teams to define API behavior and orchestrate safe, staged product launches
- Build agent workflow capabilities including task delegation, context sharing, and parallel execution patterns
- Improve reliability of long-running requests by handling timeouts, cancellation, streaming, and background execution
- Optimize request processing performance and reduce tail latency through profiling and efficient systems code
- Convert developer feedback and production incidents into improved observability, diagnostics, and product refinements
What they're looking for
- Backend services architecture and production operations
- Distributed systems and concurrency patterns
- Asynchronous execution and workflow orchestration
- Production troubleshooting using observability and profiling tools
- API design with developer empathy
- Rust (preferred)
- Streaming protocols or WebSocket implementation
- Cross-functional collaboration and technical communication
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
OpenAI
OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.
View all jobs at OpenAILikely interview questions
- Describe a time you diagnosed and fixed a production performance issue in a distributed system—what observability tools did you use?
- How would you design an API to support long-running agent workflows with reliable request handling across timeouts and cancellations?