Skip to main content

SpaceX

Site Reliability Engineer, AI Infrastructure (Starshield)

Redmond, WA$125k–$165kmidAdded today

About this role

SpaceX seeks a Site Reliability Engineer to design, deploy, and operate GPU and CPU infrastructure supporting Starshield's national security satellite constellation. You will manage large-scale AI clusters, automate Kubernetes deployments, and collaborate with engineers to ensure high-availability systems serving critical government missions.

What you'll do

  • Manage GPU/CPU infrastructure deployments to Top Secret datacenters
  • Design and productize solutions for AI clusters at 100k+ GPU scale
  • Develop automation for on-premise Kubernetes and AI cluster deployments using infrastructure-as-code tools
  • Deploy and operate core infrastructure including databases, monitoring systems, and distributed storage
  • Implement monitoring, alerting, and high-availability solutions to support critical national security systems
  • Collaborate with AI engineers throughout the full service lifecycle from design through operation and optimization

What they're looking for

  • Kubernetes cluster management and administration
  • Infrastructure-as-code tools (Terraform, Ansible)
  • Linux systems administration and boot process knowledge
  • Python, C++, or Go development
  • Bash scripting and automation
  • Containerization and OCI technologies
  • GPU deployment stacks (NVIDIA Blackwell/Rubin experience preferred)
  • Distributed systems and database management

Benefits

  • Competitive base salary ($125k–$200k depending on level)
  • Work on cutting-edge national security infrastructure
  • Opportunity to manage massive-scale AI compute systems
  • Collaboration with top-tier engineering teams
  • Global travel and operational variety
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

SpaceX

SpaceX develops advanced spacecraft and satellite systems, including the Starshield government satellite constellation and Starfall re-entry cargo capsule for global delivery. The company is hiring engineers in avionics integration, software test automation, mechanical design, and hardware reliability to validate flight-critical systems and ensure mission success.

Website
spacex.com
View all jobs at SpaceX

Likely interview questions

  • Describe your experience managing Kubernetes clusters in production—what challenges have you faced and how did you resolve them?
  • Walk us through a time you designed infrastructure automation that scaled to thousands of servers; what tools and approaches did you use?