Skip to main content

ThinkingMachines

Research Engineer, Infrastructure, Training Systems

  • Confirmed live in the last 24 hours
  • $350k–$475k
  • Mid level
  • Full-time
  • On-site · San Francisco
  • Added 2 months ago

About this role

Thinking Machines is seeking a Research Engineer to build and optimize the infrastructure powering the training of large-scale AI models. You'll design, implement, and improve distributed systems, focusing on scalability, efficiency, and reliability to empower research teams. This evergreen role welcomes candidates with a passion for scalable AI and a desire to contribute to a collaborative environment.

What you'll do

  • Design, implement, and optimize distributed training systems.
  • Develop high-performance optimizations for throughput and efficiency.
  • Create reusable frameworks and libraries for training reproducibility.
  • Establish standards for reliability, maintainability, and security.
  • Collaborate with researchers and engineers.
  • Document and share learnings through documentation, open-source contributions, or technical reports.

What they're looking for

  • Distributed training systems
  • Deep learning frameworks (PyTorch, JAX)
  • Performance optimization
  • Scalability
  • Reliability
  • System architectures
  • Computer Science

Benefits

  • Health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support
  • Visa sponsorship
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

ThinkingMachines

View all jobs at ThinkingMachines

Likely interview questions

  • Describe your experience with distributed training systems and the challenges you've faced.
  • Explain your approach to optimizing training throughput and efficiency for large models.