Anduril Industries
Site Reliability Engineer
About this role
Anduril Industries seeks a Site Reliability Engineer to own the health and uptime of fielded imaging systems deployed in real-world defense environments. You'll diagnose and resolve production issues across the full stack, manage escalations, build diagnostic tooling and runbooks, and work closely with a small team to keep systems operational in the field.
What you'll do
- Own fielded system reliability and uptime for deployed imaging systems
- Triage, diagnose, and drive resolution for production issues across networking, hardware, calibration, and software
- Serve as escalation point for customer support issues from field personnel and support channels
- Build and maintain runbooks, diagnostics, and self-service tooling to reduce recurring problems
- Reproduce software defects and hand off clean documentation to Mission Software Engineers
- Travel approximately 15% for field support and deployment windows
What they're looking for
- Linux system administration and troubleshooting
- Networking diagnostics (IP, routing, VPNs, constrained environments)
- Cross-system issue diagnosis without full component visibility
- On-call rotation and after-hours support
- Technical documentation and runbook creation
- Remote troubleshooting and non-technical communication
- Field systems and deployed hardware/software support
- Production incident response
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Anduril Industries
Anduril Industries builds autonomous defense systems including underwater vehicles, unmanned aircraft, and electronic warfare platforms for the Department of Defense. The company is hiring across mechanical engineering, mission operations, software development, technical leadership, and advanced manufacturing roles to support the design, deployment, and production of these mission-critical systems.
- Website
- anduril.com
Likely interview questions
- Walk us through a time you diagnosed a production issue that spanned multiple system layers (networking, hardware, software) with limited visibility into one component—how did you approach it?
- Describe your experience managing on-call rotations and handling after-hours escalations. How do you balance urgency with thorough root-cause analysis?