Together AI
Research Engineer, Large-Scale Training
San FranciscomidAdded today
About this role
Together AI seeks a Research Engineer to optimize large-scale foundation model training infrastructure. You'll profile and improve training systems, integrate new models, and translate cutting-edge research into production-quality solutions that power customer fine-tuning workloads.
What you'll do
- Design and optimize core components of distributed training infrastructure
- Integrate new model architectures, validate training correctness, and optimize for production workloads
- Profile distributed training to identify and eliminate bottlenecks across compute, memory, and communication
- Design experiments to validate performance hypotheses and benchmark against state-of-the-art methods
- Collaborate with Research Scientists to productionize novel training techniques
- Build and maintain experimental infrastructure balancing research velocity with production reliability
What they're looking for
- Python and PyTorch proficiency
- Multi-GPU and multi-node distributed training experience
- GPU architecture and mixed-precision training knowledge
- Distributed training paradigms (data, tensor, pipeline, expert parallelism)
- Performance profiling and systems optimization
- ML systems fundamentals
- Problem-solving from investigation through deployment
- Collaboration and technical communication
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Together AI
Together AI builds GPU compute infrastructure and open-source model customization platforms for AI developers and enterprises. The company is hiring for infrastructure operations, ML systems engineering, go-to-market technology, customer success, and GPU research roles.
- Website
- together.ai
Likely interview questions
- Describe a time you identified and resolved a performance bottleneck in a distributed training system—what was your investigation process?
- How have you approached profiling and optimizing large neural network training across multiple GPUs or nodes?