Skip to main content

Specter

Platform Site Reliability Engineer

  • Confirmed live in the last 24 hours
  • No salary listed
  • Mid level
  • Full-time
  • On-site · San Francisco
  • Added 1 month ago

About this role

Specter is seeking a Platform Site Reliability Engineer to ensure the reliability and scalability of their cloud platform powering their connected sensor fleet. This role involves maintaining Kubernetes infrastructure, managing AWS resources with Terraform, and enhancing observability and incident response. You'll collaborate with various teams to improve systems and prevent future incidents.

What you'll do

  • Debug production issues across Kubernetes, Linux systems, AWS, and applications.
  • Lead incident response and coordinate across teams during outages.
  • Manage production infrastructure with Terraform, focusing on automation and safe workflows.
  • Improve observability through logging, metrics, tracing, and alerting.
  • Define and track service-level indicators and objectives.
  • Develop runbooks and operational standards.

What they're looking for

  • Kubernetes
  • Terraform
  • AWS (EKS, IAM, networking, compute, storage)
  • Linux systems
  • Python/Go/Bash
  • Networking (DNS, load balancing, firewalls)
  • Incident Response
  • Observability
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Specter

View all jobs at Specter

Likely interview questions

  • Describe a time you had to troubleshoot a complex production issue. What was your approach?
  • How do you ensure observability and monitoring for critical systems?