OpenAI
Training Performance Engineer
About this role
OpenAI seeks a Training Performance Engineer to optimize distributed machine learning training systems. You'll profile large-scale training runs, identify bottlenecks, and implement efficiency improvements across compute, communication, and storage infrastructure to maximize cluster utilization.
What you'll do
- Profile end-to-end training runs to identify performance bottlenecks across compute, communication, and storage
- Optimize GPU utilization and throughput for large-scale distributed model training
- Collaborate with runtime and systems engineers to improve kernel efficiency and collective communication performance
- Implement model graph transforms to improve end-to-end throughput
- Build tooling to monitor and visualize MFU, throughput, and uptime metrics across clusters
- Partner with researchers to ensure new model architectures scale efficiently during pre-training
What they're looking for
- Python and C++ programming
- Distributed training on multi-GPU systems or HPC clusters
- GPU performance profiling and optimization
- PyTorch, JAX, or TensorFlow frameworks
- Debugging complex distributed systems
- CUDA (preferred)
- NCCL, MPI, or UCX communication libraries (preferred)
- Data loading and checkpointing systems (preferred)
Benefits
- Hybrid work model (3 days in office per week)
- San Francisco location
- Relocation assistance provided
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
OpenAI
OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.
View all jobs at OpenAILikely interview questions
- Walk us through how you'd approach profiling a large-scale distributed training run to identify whether bottlenecks are in compute, communication, or storage.
- Describe your experience optimizing GPU utilization and throughput on multi-GPU systems. What tools or metrics did you use to measure improvement?