Skip to main content

ThinkingMachines

Research Engineer, Infrastructure, Inference

  • Confirmed live in the last 24 hours
  • $350k–$475k
  • Mid level
  • Full-time
  • On-site · San Francisco
  • Added 2 months ago

About this role

Thinking Machines is seeking an Infrastructure Research Engineer to build and optimize the systems that power their large AI models, focusing on efficient and scalable inference. You’ll collaborate with researchers to improve performance, reliability, and reproducibility while contributing to the broader field of AI infrastructure. This is an evergreen role with ongoing evaluation of applications.

What you'll do

  • Collaborate with researchers to enable high-performance inference.
  • Design and implement techniques to improve performance, latency, throughput, and efficiency.
  • Optimize codebases and compute fleets (GPUs).
  • Extend orchestration frameworks like Kubernetes, Ray, and SLURM.
  • Establish standards for reliability and reproducibility.
  • Share learnings through documentation and potentially open-source contributions.

What they're looking for

  • Deep learning frameworks (PyTorch, JAX)
  • Inference serving systems (SGLang, vLLM)
  • Kubernetes
  • Ray
  • SLURM
  • GPU parallelism
  • Distributed compute systems

Benefits

  • Health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support
  • Visa sponsorship
  • Generous compensation
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

ThinkingMachines

View all jobs at ThinkingMachines

Likely interview questions

  • Describe your experience with optimizing inference serving systems for low latency and high throughput.
  • Explain your familiarity with Kubernetes, Ray, or SLURM – can you give an example of how you’ve used them in a distributed computing context?