Skip to main content

SpaceX

Site Reliability Engineer, AI Infrastructure (Starshield)

Washington, DC$125k–$160kmidAdded today

About this role

SpaceX seeks a Site Reliability Engineer to design, deploy, and operate GPU/CPU infrastructure supporting Starshield, a critical national security satellite constellation. You'll manage large-scale AI clusters, automate on-premise Kubernetes deployments, and ensure high availability of core infrastructure serving Top Secret datacenters.

What you'll do

  • Manage GPU/CPU infrastructure deployments to Top Secret datacenters and provide GPU-as-a-service support
  • Design and productize solutions for AI clusters scaling to 100k+ GPUs
  • Develop automation for on-premise Kubernetes, AI clusters, and operating system deployments
  • Deploy and operate core infrastructure including databases, monitoring, and distributed storage
  • Collaborate with AI engineers to build scalable, maintainable products with high availability
  • Monitor, alert, and improve system reliability throughout the service lifecycle

What they're looking for

  • Linux operating systems administration
  • Infrastructure-as-code tools (Terraform, Ansible)
  • Kubernetes cluster management and operations
  • Container technologies (OCI, Docker)
  • Python, Bash, C++, or Go scripting and development
  • GPU deployment and management (NVIDIA Blackwell/Rubin preferred)
  • Distributed databases and data modeling
  • TCP/IP networking and system performance optimization

Benefits

  • Competitive salary commensurate with experience level
  • Work on critical national security missions
  • Opportunity to scale infrastructure to unprecedented scales
  • Collaborative environment with world-class engineers
  • Global impact and innovation in AI infrastructure
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

SpaceX

SpaceX develops advanced spacecraft and satellite systems, including the Starshield government satellite constellation and Starfall re-entry cargo capsule for global delivery. The company is hiring engineers in avionics integration, software test automation, mechanical design, and hardware reliability to validate flight-critical systems and ensure mission success.

Website
spacex.com
View all jobs at SpaceX

Likely interview questions

  • Describe your experience managing Kubernetes clusters at scale—what challenges have you faced and how did you resolve them?
  • Walk us through a time you designed infrastructure automation that significantly reduced deployment time or operational overhead.