Specter
Fleet Reliability Engineer
- Confirmed live in the last 24 hours
- No salary listed
- Mid level
- Full-time
- On-site · San Francisco
- Added 2 months ago
About this role
Specter is seeking a Fleet Reliability Engineer to ensure their growing fleet of AI-powered video sensors operates reliably in the field. This role focuses on building data pipelines, analyzing telemetry, and implementing automated recovery mechanisms to proactively prevent failures and optimize fleet health. You'll translate field data into actionable insights and prioritize reliability improvements based on cost-effectiveness.
What you'll do
- Own and manage the end-to-end reliability data pipeline.
- Instrument the fleet with health metrics and optimize observability costs.
- Verify that software fixes resolve issues across the entire fleet.
- Analyze fleet telemetry to identify failure trends and root causes.
- Prioritize reliability improvements based on cost-benefit analysis.
- Establish and track fleet reliability targets (uptime, offline rate, etc.).
What they're looking for
- Python/Go
- SQL
- Observability stacks (e.g., OpenTelemetry, Grafana, Prometheus, Datadog)
- Data modeling (PostgreSQL)
- Infrastructure-as-code (Terraform)
- Telemetry analysis
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Specter
Likely interview questions
- Describe your experience building and maintaining a data pipeline end-to-end.
- How have you optimized observability costs in a previous role?