SpaceX
Site Reliability Engineer (Manufacturing Infrastructure)
About this role
SpaceX seeks a Site Reliability Engineer to own the compute, storage, and networking infrastructure supporting manufacturing operations for Starship, Starlink, Starshield, and Terafab. You'll deploy and scale systems that directly impact factory uptime and production across multiple programs, combining software engineering fundamentals with infrastructure expertise in reliability and operational excellence.
What you'll do
- Deploy, upgrade, operate, and scale compute, storage, and networking infrastructure across manufacturing systems
- Implement infrastructure as code and observability tools to ensure platform health and reliability
- Design systems for reliability and stability; identify and eliminate performance bottlenecks through measurement
- Conduct proactive maintenance including capacity planning, lifecycle management, and incident prevention
- Collaborate with software engineers and manufacturing teams to build operable, maintainable systems
- Participate in on-call rotation and travel to sites for deployments, troubleshooting, and cross-site reliability work
What they're looking for
- Linux operating systems administration
- Software development (1+ years)
- Infrastructure as code (Terraform, Ansible, Puppet)
- Containers and orchestration (Docker, Kubernetes, vSphere, KVM)
- Database management (Postgres, Clickhouse)
- Production infrastructure operations (compute, storage, networking)
- Incident response and postmortem practices
- Technical communication with diverse stakeholders
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
SpaceX
SpaceX develops advanced spacecraft and satellite systems, including the Starshield government satellite constellation and Starfall re-entry cargo capsule for global delivery. The company is hiring engineers in avionics integration, software test automation, mechanical design, and hardware reliability to validate flight-critical systems and ensure mission success.
- Website
- spacex.com
Likely interview questions
- Describe your experience managing production infrastructure at scale—what systems have you been responsible for and how did you ensure reliability?
- Walk us through how you've designed or improved infrastructure as code; what tools did you use and what was the impact on your team's efficiency?