OpenAI
Software Engineer, API Multimodal
About this role
Join OpenAI's API Multimodal team to design and operate backend services and distributed systems powering image, audio, and real-time APIs. You'll partner with research and inference teams to bring frontier multimodal capabilities to developers while owning projects from design through production at scale.
What you'll do
- Design and ship developer-facing APIs and backend services for frontier multimodal models
- Build low-latency streaming and request handling systems that reliably serve complex multimodal interactions
- Collaborate with Research teams to integrate new model capabilities into production and gather developer feedback
- Own availability, latency, scalability, and cost efficiency of deployed services
- Lead technically complex projects from ambiguous requirements through launch and iteration
- Raise engineering standards and drive architectural improvements across the team
What they're looking for
- Backend engineering (Python, Go, Rust, or TypeScript)
- Distributed systems design and architecture
- Production API and backend service development
- Low-latency streaming and real-time systems
- Observability, monitoring, and operational excellence
- Developer empathy and API design
- Cross-team collaboration and communication
- Experience with audio, image, or multimodal systems (preferred)
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
OpenAI
OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.
View all jobs at OpenAILikely interview questions
- Describe a complex backend system you designed from scratch—how did you approach latency, scalability, and cost tradeoffs?
- How would you approach integrating a new, experimental model capability into a production API while maintaining reliability?