Anduril Industries
Fielded Site Reliability Engineer
About this role
Anduril Industries seeks a Site Reliability Engineer to own the health and uptime of fielded imaging systems deployed in defense applications. You'll diagnose complex multi-stack failures, manage escalations from field personnel and customers, and transform recurring issues into durable runbooks and diagnostics.
What you'll do
- Triage and resolve fielded system outages across the full stack (networking, hardware, firmware, services)
- Serve as primary escalation point for customer support and field deployment issues
- Build and maintain runbooks, diagnostics, and self-service tooling to reduce recurring problems
- Reproduce and document software defects before handing off to product engineering teams
- Provide remote troubleshooting support to field operators and customer personnel
- Travel approximately 15% for field support and deployment windows
What they're looking for
- Linux systems administration and troubleshooting
- Networking diagnostics (IP, routing, VPNs, constrained field environments)
- Multi-component system diagnosis across hardware, firmware, and software boundaries
- Runbook and documentation creation
- Remote troubleshooting and communication with non-technical operators
- On-call rotation management and incident response
- Production support for deployed hardware/software systems
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Anduril Industries
Anduril Industries builds autonomous defense systems including underwater vehicles, unmanned aircraft, and electronic warfare platforms for the Department of Defense. The company is hiring across mechanical engineering, mission operations, software development, technical leadership, and advanced manufacturing roles to support the design, deployment, and production of these mission-critical systems.
- Website
- anduril.com
Likely interview questions
- Describe a time you diagnosed an issue in a deployed system where the root cause spanned multiple layers (hardware, networking, or software). How did you approach the troubleshooting?
- Walk us through how you'd handle a critical outage at a remote field site where you have limited direct access and must guide a non-technical operator through diagnostics.