Hadrian
Site Reliability Engineer, Robotics
About this role
Hadrian seeks a Site Reliability Engineer to ensure the stability and performance of robotics systems powering autonomous manufacturing facilities. You'll design observability infrastructure, build reliability tools, and partner across teams to embed production-grade practices into advanced manufacturing systems.
What you'll do
- Ensure reliability of robotics systems across PLCs, ROS2 middleware, and Kubernetes infrastructure
- Build observability interfaces using Prometheus, Telegraf, OpenTelemetry, and Datadog to ingest system telemetry
- Develop frameworks, diagnostic tools, and shared libraries for controls and robotics systems
- Define SLOs/SLIs and establish reliability gates with controls, robotics, and platform teams
- Create automated remediation and self-healing systems to minimize manual intervention
- Lead incident response and post-mortem analysis for production manufacturing systems
What they're looking for
- Kubernetes and container orchestration
- Infrastructure as Code and GitOps workflows
- Programming in Python, Go, TypeScript, or C++
- Systems observability and monitoring tools
- Linux fundamentals and edge infrastructure management
- ROS/ROS2 or robotics control systems experience
- Networking and on-premises deployment knowledge
- Incident management and reliability engineering
Benefits
- Medical, dental, vision, and life insurance
- 401(k) retirement plan
- Equity stake in the company
- Flexible vacation policy
- Relocation support in certain situations
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Hadrian
Hadrian builds aerospace and defense manufacturing systems, offering enterprise software platforms, advanced tooling design, and highly automated production capabilities for the sector. The company is hiring full stack engineers, manufacturing and tooling specialists, infrastructure and identity management experts, and workforce systems architects to support its rapid scaling.
View all jobs at HadrianLikely interview questions
- Tell us about a time you owned reliability for a production system where downtime had physical or operational consequences. What metrics did you track and how did you reduce incidents?
- Walk us through how you've designed observability systems for complex infrastructure. What tools have you used (Prometheus, Datadog, etc.) and how did you decide what to instrument?