Skip to main content

Together AI

Research Engineer, Large-Scale Training

San FranciscomidAdded today

About this role

Together AI seeks a Research Engineer to optimize large-scale foundation model training infrastructure. You'll profile and improve training systems, integrate new models, and translate cutting-edge research into production-quality solutions that power customer fine-tuning workloads.

What you'll do

  • Design and optimize core components of distributed training infrastructure
  • Integrate new model architectures, validate training correctness, and optimize for production workloads
  • Profile distributed training to identify and eliminate bottlenecks across compute, memory, and communication
  • Design experiments to validate performance hypotheses and benchmark against state-of-the-art methods
  • Collaborate with Research Scientists to productionize novel training techniques
  • Build and maintain experimental infrastructure balancing research velocity with production reliability

What they're looking for

  • Python and PyTorch proficiency
  • Multi-GPU and multi-node distributed training experience
  • GPU architecture and mixed-precision training knowledge
  • Distributed training paradigms (data, tensor, pipeline, expert parallelism)
  • Performance profiling and systems optimization
  • ML systems fundamentals
  • Problem-solving from investigation through deployment
  • Collaboration and technical communication
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Together AI

Together AI builds GPU compute infrastructure and open-source model customization platforms for AI developers and enterprises. The company is hiring for infrastructure operations, ML systems engineering, go-to-market technology, customer success, and GPU research roles.

View all jobs at Together AI

Likely interview questions

  • Describe a time you identified and resolved a performance bottleneck in a distributed training system—what was your investigation process?
  • How have you approached profiling and optimizing large neural network training across multiple GPUs or nodes?