Picogrid
Site Reliability Engineer
About this role
Picogrid seeks their first Site Reliability Engineer to own production reliability across cloud and edge infrastructure supporting defense operations worldwide. You'll design SLOs, manage observability stacks, lead incident response, and build reliability practices for systems deployed in challenging battlefield conditions.
What you'll do
- Define and drive SLI/SLO targets for cloud and remote edge device deployments
- Manage observability infrastructure including Grafana, Prometheus, Loki, and OpenTelemetry with versioned dashboards and alert rules
- Participate in on-call rotation, incident triage, postmortems, and infrastructure hardening
- Encode reliability practices into infrastructure-as-code deployments
- Oversee Kubernetes operations including node lifecycle, workload scheduling, and cluster debugging
- Support high-availability database and edge fleet management across contested environments
What they're looking for
- Kubernetes operations (node lifecycle, StatefulSets, graceful drains, debugging)
- Observability design (Grafana, Prometheus, Loki, OpenTelemetry)
- Infrastructure-as-code (Terraform or OpenTofu)
- AWS (IAM, networking, multi-account, hardening)
- Incident response and blameless postmortem practices
- IoT and edge fleet operations
- High-availability database deployments
- Overlay/mesh networking (Nebula, WireGuard, Tailscale)
Benefits
- Stock options with early-stage upside potential
- 401(k) with employer matching
- Full health coverage (medical, dental, vision)
- Unlimited PTO (minimum two weeks) plus 11 paid holidays
- Paid parental leave and relocation assistance
- In-office perks: free lunch, EV charging, premium coffee and snacks
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Picogrid
Picogrid builds advanced defense technology systems including autonomous platforms and complex software solutions for military applications. The company is hiring Field Operations Engineers, Fullstack Engineers, and Forward Deployed Engineers to support product installation, development, and direct customer delivery.
View all jobs at PicogridLikely interview questions
- Describe your most complex production incident—what was your investigation process and how did you prevent recurrence?
- How have you approached designing SLOs and alerting thresholds to balance reliability with alert fatigue?