Skip to main content

ThinkingMachines

Research Engineer, Infrastructure, Kernels

  • Confirmed live in the last 24 hours
  • $350k–$475k
  • Mid level
  • Full-time
  • On-site · San Francisco
  • Added 2 months ago

About this role

Thinking Machines is seeking a Research Engineer to build and optimize the infrastructure powering their large-scale language model training. You'll focus on designing high-performance ML kernels, improving compute efficiency, and collaborating with research teams to scale AI systems. This is an ongoing, 'evergreen' role, so express your interest even if your experience isn't a perfect match.

What you'll do

  • Design and implement custom ML kernels (CUDA, CuTe, Triton) for LLM operations.
  • Optimize compute primitives to reduce memory bandwidth bottlenecks.
  • Collaborate with research teams to align kernel optimizations with model architecture.
  • Develop and maintain a library of reusable kernels and performance benchmarks.
  • Contribute to infrastructure stability and scalability.
  • Document and share insights through internal talks and technical papers.

What they're looking for

  • CUDA
  • CuTe
  • Triton
  • PyTorch
  • JAX
  • Deep Learning Frameworks
  • GPU Programming
  • Compute Optimization

Benefits

  • Health, Dental, and Vision Benefits
  • Unlimited PTO
  • Paid Parental Leave
  • Relocation Support
  • Visa Sponsorship
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

ThinkingMachines

View all jobs at ThinkingMachines

Likely interview questions

  • Describe a time you optimized a compute-intensive workload. What were the key challenges and how did you overcome them?
  • Explain your experience with CUDA, CuTe, or Triton. Can you describe a specific kernel you developed or modified?