Skip to main content

ThinkingMachines

Site Reliability Engineer, Post Training

  • Confirmed live in the last 24 hours
  • $300k–$350k
  • Mid level
  • Full-time
  • On-site · San Francisco
  • 4+ yrs exp
  • Added 1 month ago

About this role

Thinking Machines is seeking a Site Reliability Engineer to ensure the reliability and performance of their post-training and reinforcement learning systems. This role involves direct collaboration with research teams, debugging production issues, and building automation to streamline model training. You'll be a key point of contact for production issues and instrumental in preventing future incidents.

What you'll do

  • Ensure reliability, performance, and uptime of training jobs.
  • Collaborate with research teams during active model runs.
  • Debug failures across the full stack (accelerators, networking, storage, etc.).
  • Build monitoring, alerting, and automated recovery systems.
  • Improve checkpointing, fault tolerance, and job scheduling.
  • Build internal tools to reduce toil and improve cluster utilization.

What they're looking for

  • Python
  • Linux systems internals
  • Networking fundamentals
  • Distributed systems
  • Kubernetes
  • PyTorch
  • Ray

Benefits

  • Health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support
  • Visa sponsorship
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

ThinkingMachines

View all jobs at ThinkingMachines

Likely interview questions

  • Describe a time you debugged a complex failure in a distributed system. What was your approach?
  • How do you balance automation and human intervention when addressing production issues?