Skip to main content

Incident IQ

Site Reliability Engineer

  • Confirmed live in the last 24 hours
  • No salary listed
  • Mid level
  • Remote
  • 5+ yrs exp
  • Added 2 months ago

About this role

Incident IQ is seeking a first-of-its-kind Site Reliability Engineer to build their reliability foundation from the ground up. This hands-on role focuses on defining reliability standards, implementing observability tooling like Grafana and PagerDuty, and automating operational processes to ensure the stability and performance of their K-12 workflow platform. The ideal candidate thrives in a fast-paced startup environment, proactively identifies and solves problems, and effectively communicates technical concepts.

What you'll do

  • Define and implement Service Level Indicators (SLIs) and Service Level Objectives (SLOs).
  • Establish and continuously improve the incident management practice.
  • Own and manage the observability stack (metrics, logs, traces, RUM, synthetic checks).
  • Partner with engineering teams to refine SLIs, SLOs, and error budgets.
  • Automate repetitive operational tasks through infrastructure as code.
  • Design and conduct load/performance tests and chaos engineering exercises.

What they're looking for

  • Site Reliability Engineering (SRE)
  • Observability (metrics, logs, traces)
  • Grafana
  • PagerDuty
  • Infrastructure as Code
  • Incident Management
  • SLI/SLO Methodology
  • Operating Systems
  • Networking
  • AI Tools
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Incident IQ

View all jobs at Incident IQ

Likely interview questions

  • Describe a time you defined and implemented an SLO. What were the challenges?
  • How do you approach incident management and post-incident analysis?