ThinkingMachines
Research Engineer, Infrastructure, Training Systems
- Confirmed live in the last 24 hours
- $350k–$475k
- Mid level
- Full-time
- On-site · San Francisco
- Added 2 months ago
About this role
Thinking Machines is seeking a Research Engineer to build and optimize the infrastructure powering the training of large-scale AI models. You'll design, implement, and improve distributed systems, focusing on scalability, efficiency, and reliability to empower research teams. This evergreen role welcomes candidates with a passion for scalable AI and a desire to contribute to a collaborative environment.
What you'll do
- Design, implement, and optimize distributed training systems.
- Develop high-performance optimizations for throughput and efficiency.
- Create reusable frameworks and libraries for training reproducibility.
- Establish standards for reliability, maintainability, and security.
- Collaborate with researchers and engineers.
- Document and share learnings through documentation, open-source contributions, or technical reports.
What they're looking for
- Distributed training systems
- Deep learning frameworks (PyTorch, JAX)
- Performance optimization
- Scalability
- Reliability
- System architectures
- Computer Science
Benefits
- Health, dental, and vision benefits
- Unlimited PTO
- Paid parental leave
- Relocation support
- Visa sponsorship
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
ThinkingMachines
Likely interview questions
- Describe your experience with distributed training systems and the challenges you've faced.
- Explain your approach to optimizing training throughput and efficiency for large models.