Cerebras
Cluster Operations Software Engineer
- Confirmed live in the last 24 hours
- No salary listed
- Mid level
- Full-time
- Remote · Sunnyvale, CA
- 6+ yrs exp
- Added 1 month ago
About this role
Cerebras Systems is seeking a Cluster Operations Software Engineer to maintain and optimize their cutting-edge AI compute clusters, powered by the world's largest AI chip, the Wafer-Scale Engine (WSE). This role is essential for ensuring infrastructure health, performance, and availability, and involves building tools and automation to support large-scale AI initiatives. The ideal candidate will thrive in a fast-paced environment and possess a strong background in distributed systems and Linux-based infrastructure.
What you'll do
- Deploy, configure, and debug container-based services with Docker.
- Develop and maintain operational dashboards and reliability tooling.
- Collaborate with cross-functional teams to improve operational visibility.
- Manage and operate multiple AI compute infrastructure clusters.
- Monitor and proactively resolve cluster health issues.
- Maximize compute capacity through optimization.
What they're looking for
- Python
- Docker
- Linux
- Distributed Systems
- Monitoring and Alerting Systems
- Troubleshooting
Benefits
- Work on the world's fastest AI supercomputer
- Startup vitality with job stability
- Non-corporate work culture
- 24/7 on-call rotation
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Cerebras
Likely interview questions
- Describe your experience managing and operating large-scale compute infrastructure, specifically in the context of machine learning or HPC.
- Explain your experience with Docker and Kubernetes. How have you used these tools to automate deployments or manage containers?