ThinkingMachines
Site Reliability Engineer, Post Training
- Confirmed live in the last 24 hours
- $300k–$350k
- Mid level
- Full-time
- On-site · San Francisco
- 4+ yrs exp
- Added 1 month ago
About this role
Thinking Machines is seeking a Site Reliability Engineer to ensure the reliability and performance of their post-training and reinforcement learning systems. This role involves direct collaboration with research teams, debugging production issues, and building automation to streamline model training. You'll be a key point of contact for production issues and instrumental in preventing future incidents.
What you'll do
- Ensure reliability, performance, and uptime of training jobs.
- Collaborate with research teams during active model runs.
- Debug failures across the full stack (accelerators, networking, storage, etc.).
- Build monitoring, alerting, and automated recovery systems.
- Improve checkpointing, fault tolerance, and job scheduling.
- Build internal tools to reduce toil and improve cluster utilization.
What they're looking for
- Python
- Linux systems internals
- Networking fundamentals
- Distributed systems
- Kubernetes
- PyTorch
- Ray
Benefits
- Health, dental, and vision benefits
- Unlimited PTO
- Paid parental leave
- Relocation support
- Visa sponsorship
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
ThinkingMachines
Likely interview questions
- Describe a time you debugged a complex failure in a distributed system. What was your approach?
- How do you balance automation and human intervention when addressing production issues?