ThinkingMachines
Research Engineer, Infrastructure, Numerics
- Confirmed live in the last 24 hours
- $350k–$475k
- Mid level
- Full-time
- On-site · San Francisco
- Added 2 months ago
About this role
Thinking Machines is seeking a Research Engineer to build and optimize infrastructure for training large-scale AI models, with a focus on numerical precision and performance. This role sits at the intersection of research and systems engineering, requiring a deep understanding of both mathematical optimization and distributed computing. You'll collaborate closely with research teams to ensure stability, scalability, and efficiency in training trillion-parameter models.
What you'll do
- Design and optimize distributed training infrastructure for LLMs.
- Implement and evaluate low-precision numerics (e.g., BF16, MXFP8, NVFP4).
- Develop optimized kernels and communication primitives.
- Collaborate with research teams on model architectures and training strategies.
- Prototype and benchmark scaling strategies.
- Contribute to orchestration and monitoring systems.
What they're looking for
- Deep Learning Frameworks (PyTorch, JAX)
- Distributed Systems
- Floating-point Numerics
- Low-precision Arithmetic
- PyTorch/XLA
- DeepSpeed
- Megatron-LM
- FP8/INT8/MX formats
Benefits
- Health, dental, and vision benefits
- Unlimited PTO
- Paid parental leave
- Relocation support
- Visa sponsorship
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
ThinkingMachines
Likely interview questions
- Describe your experience with distributed training frameworks like PyTorch or JAX.
- Explain your understanding of low-precision numerics and their trade-offs.