Varda Space Industries
Site Reliability Engineer II
About this role
Varda Space Industries seeks an experienced Site Reliability Engineer II to build and maintain mission-critical infrastructure supporting spacecraft and Earth-based systems. You'll apply software engineering principles to containerized environments, infrastructure automation, and production reliability while working in a fast-paced aerospace technology company.
What you'll do
- Deploy, maintain, and operate mission-critical applications and infrastructure for spacecraft and company systems
- Build and evolve Infrastructure as Code using Terraform and manage Kubernetes clusters in production
- Implement observability systems (metrics, logging, tracing) with actionable alerting
- Design and maintain CI/CD pipelines for safe, repeatable deployments
- Respond to production incidents, perform root cause analysis, and lead blameless postmortems
- Partner with software and hardware teams to optimize system reliability, scalability, and developer experience
What they're looking for
- Kubernetes and container orchestration
- Infrastructure as Code (Terraform)
- Prometheus, Grafana, or similar observability tools
- Python, Bash, or PowerShell scripting
- Software-defined networking (VPCs, firewalls, VPNs)
- CI/CD pipeline design and implementation
- Azure cloud infrastructure (preferred)
- Configuration management and GitOps tools
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Varda Space Industries
Varda Space Industries builds commercial spacecraft and orbital manufacturing systems for missions in low Earth orbit, including reentry capsules and satellite buses. The company is hiring mission operations engineers, software engineers (ground segment, flight, and embedded), and manufacturing engineers to develop mission-critical systems, operational infrastructure, and scalable production capabilities.
- Website
- varda.com
Likely interview questions
- Walk us through a complex production incident you resolved—what was your troubleshooting approach and how did you prevent recurrence?
- Describe your experience operating Kubernetes in production; what challenges did you face and how did you overcome them?