OpenAI
Machine Learning Engineer, API Multicloud
About this role
OpenAI is seeking a Machine Learning Engineer to build and improve AI systems enabling strategic partners to adapt OpenAI models in cloud-native AWS environments. You'll work across post-training workflows, evaluation, data pipelines, and infrastructure integration, partnering with customers and internal teams to solve complex model-performance problems and scale production ML systems.
What you'll do
- Partner with strategic customers and internal teams to define target model behaviors and translate needs into training and evaluation requirements
- Build and scale production ML systems for model customization, post-training, and fine-tuning-as-a-service workflows
- Investigate training and customization workflows to identify improvements to data, evaluation, training, or infrastructure
- Integrate ML capabilities into AWS-native API environments in collaboration with backend and infrastructure engineers
- Propose and implement improvements to post-training systems, tooling, APIs, and developer workflows based on partner feedback
- Debug and improve complex systems spanning model behavior, training data, APIs, distributed infrastructure, and product surfaces
What they're looking for
- Deep learning and transformer models (PyTorch or TensorFlow)
- Large language model training, fine-tuning, and post-training techniques
- Production ML systems design and deployment
- Data pipelines and evaluation systems
- Python or Rust production code
- Distributed systems and cloud infrastructure (AWS preferred)
- Systems design and software engineering fundamentals
- Model customization and behavior analysis
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
OpenAI
OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.
View all jobs at OpenAILikely interview questions
- Walk us through a time you debugged a complex production ML system spanning multiple components—how did you isolate the root cause?
- Describe your experience with fine-tuning or post-training large language models. What techniques have you implemented and what challenges did you face?