SpaceX
Site Reliability Engineer, AI Infrastructure (Starshield)
About this role
SpaceX is hiring a Site Reliability Engineer to design, deploy, and operate GPU/CPU infrastructure supporting Starshield, a US government satellite constellation for national security missions. You'll manage large-scale AI clusters, automate infrastructure deployments to secure data centers, and collaborate with AI engineers to build highly available systems.
What you'll do
- Manage GPU/CPU infrastructure deployments to classified data centers
- Design and productize solutions for massive AI clusters (100k+ GPU scale)
- Develop automation for Kubernetes and AI cluster deployment using infrastructure-as-code tools
- Operate core infrastructure including databases, monitoring systems, and distributed storage
- Implement monitoring, alerting, and high-availability solutions for critical systems
- Collaborate with AI engineers throughout the service lifecycle from design to operation
What they're looking for
- Linux operating systems administration
- Infrastructure-as-code tools (Terraform, Ansible)
- Kubernetes cluster management and operations
- Python and C++ or Go development
- Bash scripting and automation
- GPU deployment and management (NVIDIA stacks)
- Distributed systems and databases
- CI/CD, testing, and monitoring practices
Benefits
- Competitive base salary ($125,000–$195,000 depending on level)
- Work on critical national security missions
- Opportunity to scale infrastructure to 100k+ GPU systems
- Collaboration with world-class AI and infrastructure engineers
- Exposure to cutting-edge satellite and AI technology
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
SpaceX
SpaceX develops advanced spacecraft and satellite systems, including the Starshield government satellite constellation and Starfall re-entry cargo capsule for global delivery. The company is hiring engineers in avionics integration, software test automation, mechanical design, and hardware reliability to validate flight-critical systems and ensure mission success.
- Website
- spacex.com
Likely interview questions
- Describe your experience managing Kubernetes clusters at scale—what challenges have you faced and how did you resolve them?
- Walk us through how you'd design automation to deploy and manage thousands of GPU servers across multiple data centers.