Skip to main content

ThinkingMachines

Research Engineer, Infrastructure, Numerics

  • Confirmed live in the last 24 hours
  • $350k–$475k
  • Mid level
  • Full-time
  • On-site · San Francisco
  • Added 2 months ago

About this role

Thinking Machines is seeking a Research Engineer to build and optimize infrastructure for training large-scale AI models, with a focus on numerical precision and performance. This role sits at the intersection of research and systems engineering, requiring a deep understanding of both mathematical optimization and distributed computing. You'll collaborate closely with research teams to ensure stability, scalability, and efficiency in training trillion-parameter models.

What you'll do

  • Design and optimize distributed training infrastructure for LLMs.
  • Implement and evaluate low-precision numerics (e.g., BF16, MXFP8, NVFP4).
  • Develop optimized kernels and communication primitives.
  • Collaborate with research teams on model architectures and training strategies.
  • Prototype and benchmark scaling strategies.
  • Contribute to orchestration and monitoring systems.

What they're looking for

  • Deep Learning Frameworks (PyTorch, JAX)
  • Distributed Systems
  • Floating-point Numerics
  • Low-precision Arithmetic
  • PyTorch/XLA
  • DeepSpeed
  • Megatron-LM
  • FP8/INT8/MX formats

Benefits

  • Health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support
  • Visa sponsorship
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

ThinkingMachines

View all jobs at ThinkingMachines

Likely interview questions

  • Describe your experience with distributed training frameworks like PyTorch or JAX.
  • Explain your understanding of low-precision numerics and their trade-offs.