PagerDuty
Site Reliability Engineer I
Atlanta$98k–$148.5kentryAdded today
About this role
PagerDuty is seeking a Site Reliability Engineer I to build and operate foundational infrastructure supporting millions of daily events and alerts. You'll work on networking, compute platforms, and Kubernetes systems while ensuring reliability and scalability for a platform trusted by Fortune 100 companies.
What you'll do
- Support and improve foundational infrastructure including networking, compute, Kubernetes clusters, and ingress/traffic management
- Enhance reliability and scalability by hardening systems and rolling out new infrastructure capabilities
- Participate in agile rituals and communicate progress and risks to the team
- Stay current on technical trends and propose innovative tools and approaches
- Monitor system health using metrics, logs, and alerts; participate in 24/7 on-call rotations
- Detect, respond to, and resolve production incidents
What they're looking for
- Linux system administration in production environments
- Container orchestration with Kubernetes or EKS
- Cloud infrastructure on AWS, GCP, or Azure
- Networking fundamentals (load balancing, DNS, TLS, ingress)
- Infrastructure as Code (Terraform, CloudFormation)
- Programming in Python, Ruby, Go, or similar languages
- Monitoring and observability platforms (DataDog, Prometheus, Grafana)
- Service meshes or ingress controllers (Envoy, Istio, NGINX)
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
PagerDuty
- Website
- pagerduty.com
Likely interview questions
- Walk us through your experience operating Linux systems in production—what was the most complex incident you debugged?
- Describe your hands-on experience with Kubernetes or EKS. What aspects of cluster operations are you most comfortable with?