ThinkingMachines
Software Engineer, Supercomputing
- Confirmed live in the last 24 hours
- $350k–$475k
- Mid level
- Full-time
- On-site · San Francisco
- Added 2 months ago
About this role
Thinking Machines is seeking a skilled Software Engineer to build and maintain their high-performance GPU supercomputing environment. This role focuses on ensuring reliable, cost-effective compute resources for AI model training and inference at scale. It's an evergreen role, meaning they are continually reviewing applications for potential future opportunities.
What you'll do
- Operating and automating large GPU clusters.
- Developing software to abstract cluster management.
- Extending scheduling/orchestration systems like Kubernetes or Slurm.
- Monitoring and improving operational metrics.
- Building reliable storage and artifact paths.
- Collaborating with researchers to optimize performance.
What they're looking for
- Python
- Rust
- Kubernetes
- Slurm
- Linux
- Networking
- Infrastructure-as-Code
- CUDA/NCCL
Benefits
- Health, dental, and vision benefits
- Unlimited PTO
- Paid parental leave
- Relocation support
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
ThinkingMachines
Likely interview questions
- Describe your experience operating large-scale GPU clusters.
- How have you used Kubernetes or Slurm to optimize resource allocation?