PagerDuty
Site Reliability Engineer II
Atlanta$113k–$171.6kmidAdded today
About this role
PagerDuty seeks a Site Reliability Engineer II to build and operate core infrastructure powering their digital operations platform. You'll manage foundational networking, compute, and Kubernetes systems that process millions of events daily, while participating in on-call rotations and driving platform reliability at scale.
What you'll do
- Design, build, and maintain foundational infrastructure including networking, compute platforms, and Kubernetes clusters
- Harden existing systems and support rollout of new infrastructure capabilities to improve reliability and scalability
- Monitor system health using metrics, logs, and alerts; respond to incidents through 24/7 on-call rotations
- Participate in agile rituals including standups, planning sessions, and retrospectives
- Evaluate and recommend innovative tools and technical approaches for infrastructure challenges
- Manage traffic ingress systems and load balancing across cloud-native environments
What they're looking for
- Site Reliability Engineering and DevOps practices
- Linux system administration in production environments
- Kubernetes and container orchestration (EKS)
- Cloud infrastructure (AWS, GCP, or Azure)
- Infrastructure as Code (Terraform, CloudFormation)
- Programming (Python, Ruby, Go, or similar)
- Networking fundamentals (DNS, TLS, load balancing, ingress)
- Monitoring and observability platforms
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
PagerDuty
- Website
- pagerduty.com
Likely interview questions
- Walk us through a production incident you debugged involving Kubernetes or cloud infrastructure—what was your approach and what did you learn?
- Describe your experience managing infrastructure-as-code pipelines and how you've versioned or rolled back changes.