SpaceX
Site Reliability Engineer, AI Infrastructure (Starshield)
About this role
SpaceX seeks a Site Reliability Engineer to design, deploy, and operate GPU and CPU infrastructure supporting Starshield's national security satellite constellation. You will manage large-scale AI clusters, automate Kubernetes deployments, and collaborate with engineers to ensure high-availability systems serving critical government missions.
What you'll do
- Manage GPU/CPU infrastructure deployments to Top Secret datacenters
- Design and productize solutions for AI clusters at 100k+ GPU scale
- Develop automation for on-premise Kubernetes and AI cluster deployments using infrastructure-as-code tools
- Deploy and operate core infrastructure including databases, monitoring systems, and distributed storage
- Implement monitoring, alerting, and high-availability solutions to support critical national security systems
- Collaborate with AI engineers throughout the full service lifecycle from design through operation and optimization
What they're looking for
- Kubernetes cluster management and administration
- Infrastructure-as-code tools (Terraform, Ansible)
- Linux systems administration and boot process knowledge
- Python, C++, or Go development
- Bash scripting and automation
- Containerization and OCI technologies
- GPU deployment stacks (NVIDIA Blackwell/Rubin experience preferred)
- Distributed systems and database management
Benefits
- Competitive base salary ($125k–$200k depending on level)
- Work on cutting-edge national security infrastructure
- Opportunity to manage massive-scale AI compute systems
- Collaboration with top-tier engineering teams
- Global travel and operational variety
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
SpaceX
SpaceX develops advanced spacecraft and satellite systems, including the Starshield government satellite constellation and Starfall re-entry cargo capsule for global delivery. The company is hiring engineers in avionics integration, software test automation, mechanical design, and hardware reliability to validate flight-critical systems and ensure mission success.
- Website
- spacex.com
Likely interview questions
- Describe your experience managing Kubernetes clusters in production—what challenges have you faced and how did you resolve them?
- Walk us through a time you designed infrastructure automation that scaled to thousands of servers; what tools and approaches did you use?