Skip to main content

OpenAI

Training Performance Engineer

San Francisco (Remote)$250k–$445kfulltimemidAdded 1 month ago

About this role

OpenAI seeks a Training Performance Engineer to optimize distributed machine learning training systems. You'll profile large-scale training runs, identify bottlenecks, and implement efficiency improvements across compute, communication, and storage infrastructure to maximize cluster utilization.

What you'll do

  • Profile end-to-end training runs to identify performance bottlenecks across compute, communication, and storage
  • Optimize GPU utilization and throughput for large-scale distributed model training
  • Collaborate with runtime and systems engineers to improve kernel efficiency and collective communication performance
  • Implement model graph transforms to improve end-to-end throughput
  • Build tooling to monitor and visualize MFU, throughput, and uptime metrics across clusters
  • Partner with researchers to ensure new model architectures scale efficiently during pre-training

What they're looking for

  • Python and C++ programming
  • Distributed training on multi-GPU systems or HPC clusters
  • GPU performance profiling and optimization
  • PyTorch, JAX, or TensorFlow frameworks
  • Debugging complex distributed systems
  • CUDA (preferred)
  • NCCL, MPI, or UCX communication libraries (preferred)
  • Data loading and checkpointing systems (preferred)

Benefits

  • Hybrid work model (3 days in office per week)
  • San Francisco location
  • Relocation assistance provided
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

OpenAI

OpenAI builds AI infrastructure and products, including large-scale data center campuses for AI computing and generative AI applications for enterprise customers. The company is hiring civil engineers, project engineers, electrical design engineers, data center R&D engineers, and AI deployment engineers to expand its infrastructure capabilities and help customers deploy AI solutions.

View all jobs at OpenAI

Likely interview questions

  • Walk us through how you'd approach profiling a large-scale distributed training run to identify whether bottlenecks are in compute, communication, or storage.
  • Describe your experience optimizing GPU utilization and throughput on multi-GPU systems. What tools or metrics did you use to measure improvement?