Incident IQ
Site Reliability Engineer
- Confirmed live in the last 24 hours
- No salary listed
- Mid level
- Remote
- 5+ yrs exp
- Added 2 months ago
About this role
Incident IQ is seeking a first-of-its-kind Site Reliability Engineer to build their reliability foundation from the ground up. This hands-on role focuses on defining reliability standards, implementing observability tooling like Grafana and PagerDuty, and automating operational processes to ensure the stability and performance of their K-12 workflow platform. The ideal candidate thrives in a fast-paced startup environment, proactively identifies and solves problems, and effectively communicates technical concepts.
What you'll do
- Define and implement Service Level Indicators (SLIs) and Service Level Objectives (SLOs).
- Establish and continuously improve the incident management practice.
- Own and manage the observability stack (metrics, logs, traces, RUM, synthetic checks).
- Partner with engineering teams to refine SLIs, SLOs, and error budgets.
- Automate repetitive operational tasks through infrastructure as code.
- Design and conduct load/performance tests and chaos engineering exercises.
What they're looking for
- Site Reliability Engineering (SRE)
- Observability (metrics, logs, traces)
- Grafana
- PagerDuty
- Infrastructure as Code
- Incident Management
- SLI/SLO Methodology
- Operating Systems
- Networking
- AI Tools
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Incident IQ
Likely interview questions
- Describe a time you defined and implemented an SLO. What were the challenges?
- How do you approach incident management and post-incident analysis?