ThinkingMachines
Research Engineer, Infrastructure, Kernels
- Confirmed live in the last 24 hours
- $350k–$475k
- Mid level
- Full-time
- On-site · San Francisco
- Added 2 months ago
About this role
Thinking Machines is seeking a Research Engineer to build and optimize the infrastructure powering their large-scale language model training. You'll focus on designing high-performance ML kernels, improving compute efficiency, and collaborating with research teams to scale AI systems. This is an ongoing, 'evergreen' role, so express your interest even if your experience isn't a perfect match.
What you'll do
- Design and implement custom ML kernels (CUDA, CuTe, Triton) for LLM operations.
- Optimize compute primitives to reduce memory bandwidth bottlenecks.
- Collaborate with research teams to align kernel optimizations with model architecture.
- Develop and maintain a library of reusable kernels and performance benchmarks.
- Contribute to infrastructure stability and scalability.
- Document and share insights through internal talks and technical papers.
What they're looking for
- CUDA
- CuTe
- Triton
- PyTorch
- JAX
- Deep Learning Frameworks
- GPU Programming
- Compute Optimization
Benefits
- Health, Dental, and Vision Benefits
- Unlimited PTO
- Paid Parental Leave
- Relocation Support
- Visa Sponsorship
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
ThinkingMachines
Likely interview questions
- Describe a time you optimized a compute-intensive workload. What were the key challenges and how did you overcome them?
- Explain your experience with CUDA, CuTe, or Triton. Can you describe a specific kernel you developed or modified?