Skip to main content

ThinkingMachines

Software Engineer, Supercomputing

  • Confirmed live in the last 24 hours
  • $350k–$475k
  • Mid level
  • Full-time
  • On-site · San Francisco
  • Added 2 months ago

About this role

Thinking Machines is seeking a skilled Software Engineer to build and maintain their high-performance GPU supercomputing environment. This role focuses on ensuring reliable, cost-effective compute resources for AI model training and inference at scale. It's an evergreen role, meaning they are continually reviewing applications for potential future opportunities.

What you'll do

  • Operating and automating large GPU clusters.
  • Developing software to abstract cluster management.
  • Extending scheduling/orchestration systems like Kubernetes or Slurm.
  • Monitoring and improving operational metrics.
  • Building reliable storage and artifact paths.
  • Collaborating with researchers to optimize performance.

What they're looking for

  • Python
  • Rust
  • Kubernetes
  • Slurm
  • Linux
  • Networking
  • Infrastructure-as-Code
  • CUDA/NCCL

Benefits

  • Health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

ThinkingMachines

View all jobs at ThinkingMachines

Likely interview questions

  • Describe your experience operating large-scale GPU clusters.
  • How have you used Kubernetes or Slurm to optimize resource allocation?