Skip to main content

Cerebras

Cluster Operations Software Engineer

  • Confirmed live in the last 24 hours
  • No salary listed
  • Mid level
  • Full-time
  • Remote · Sunnyvale, CA
  • 6+ yrs exp
  • Added 1 month ago

About this role

Cerebras Systems is seeking a Cluster Operations Software Engineer to maintain and optimize their cutting-edge AI compute clusters, powered by the world's largest AI chip, the Wafer-Scale Engine (WSE). This role is essential for ensuring infrastructure health, performance, and availability, and involves building tools and automation to support large-scale AI initiatives. The ideal candidate will thrive in a fast-paced environment and possess a strong background in distributed systems and Linux-based infrastructure.

What you'll do

  • Deploy, configure, and debug container-based services with Docker.
  • Develop and maintain operational dashboards and reliability tooling.
  • Collaborate with cross-functional teams to improve operational visibility.
  • Manage and operate multiple AI compute infrastructure clusters.
  • Monitor and proactively resolve cluster health issues.
  • Maximize compute capacity through optimization.

What they're looking for

  • Python
  • Docker
  • Linux
  • Distributed Systems
  • Monitoring and Alerting Systems
  • Troubleshooting

Benefits

  • Work on the world's fastest AI supercomputer
  • Startup vitality with job stability
  • Non-corporate work culture
  • 24/7 on-call rotation
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Cerebras

View all jobs at Cerebras

Likely interview questions

  • Describe your experience managing and operating large-scale compute infrastructure, specifically in the context of machine learning or HPC.
  • Explain your experience with Docker and Kubernetes. How have you used these tools to automate deployments or manage containers?