Skip to main content

Specter

Fleet Reliability Engineer

  • Confirmed live in the last 24 hours
  • No salary listed
  • Mid level
  • Full-time
  • On-site · San Francisco
  • Added 2 months ago

About this role

Specter is seeking a Fleet Reliability Engineer to ensure their growing fleet of AI-powered video sensors operates reliably in the field. This role focuses on building data pipelines, analyzing telemetry, and implementing automated recovery mechanisms to proactively prevent failures and optimize fleet health. You'll translate field data into actionable insights and prioritize reliability improvements based on cost-effectiveness.

What you'll do

  • Own and manage the end-to-end reliability data pipeline.
  • Instrument the fleet with health metrics and optimize observability costs.
  • Verify that software fixes resolve issues across the entire fleet.
  • Analyze fleet telemetry to identify failure trends and root causes.
  • Prioritize reliability improvements based on cost-benefit analysis.
  • Establish and track fleet reliability targets (uptime, offline rate, etc.).

What they're looking for

  • Python/Go
  • SQL
  • Observability stacks (e.g., OpenTelemetry, Grafana, Prometheus, Datadog)
  • Data modeling (PostgreSQL)
  • Infrastructure-as-code (Terraform)
  • Telemetry analysis
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Specter

View all jobs at Specter

Likely interview questions

  • Describe your experience building and maintaining a data pipeline end-to-end.
  • How have you optimized observability costs in a previous role?