ThinkingMachines
Site Reliability Engineer, Production
- Confirmed live in the last 24 hours
- $300k–$350k
- Mid level
- Full-time
- On-site · San Francisco
- Added 1 month ago
About this role
Thinking Machines is seeking a Site Reliability Engineer to ensure the robustness and resilience of their Tinker platform, a fine-tuning API for AI models. You’ll collaborate closely with engineering and research teams to define and maintain system reliability, responding to incidents and implementing preventative measures. This role focuses on optimizing Tinker's performance, security, and scalability as the platform grows.
What you'll do
- Define and own end-to-end system reliability
- Develop and implement Service Level Objectives (SLOs)
- Design and implement monitoring and observability systems
- Lead incident response and post-incident reviews
- Improve multi-tenant isolation and resource scheduling
What they're looking for
- Distributed systems
- Cloud infrastructure
- Site reliability engineering
- Incident response
- Kubernetes
- Automation
- Communication
- Software development
Benefits
- Health, dental, and vision benefits
- Unlimited PTO
- Paid parental leave
- Relocation support
- Visa sponsorship
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
ThinkingMachines
Likely interview questions
- Describe your experience with production incident response and postmortems.
- How do you approach defining and implementing SLOs for distributed systems?