Skip to main content

PagerDuty

Site Reliability Engineer I

Atlanta$98k–$148.5kentryAdded today

About this role

PagerDuty is seeking a Site Reliability Engineer I to build and operate foundational infrastructure supporting millions of daily events and alerts. You'll work on networking, compute platforms, and Kubernetes systems while ensuring reliability and scalability for a platform trusted by Fortune 100 companies.

What you'll do

  • Support and improve foundational infrastructure including networking, compute, Kubernetes clusters, and ingress/traffic management
  • Enhance reliability and scalability by hardening systems and rolling out new infrastructure capabilities
  • Participate in agile rituals and communicate progress and risks to the team
  • Stay current on technical trends and propose innovative tools and approaches
  • Monitor system health using metrics, logs, and alerts; participate in 24/7 on-call rotations
  • Detect, respond to, and resolve production incidents

What they're looking for

  • Linux system administration in production environments
  • Container orchestration with Kubernetes or EKS
  • Cloud infrastructure on AWS, GCP, or Azure
  • Networking fundamentals (load balancing, DNS, TLS, ingress)
  • Infrastructure as Code (Terraform, CloudFormation)
  • Programming in Python, Ruby, Go, or similar languages
  • Monitoring and observability platforms (DataDog, Prometheus, Grafana)
  • Service meshes or ingress controllers (Envoy, Istio, NGINX)
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

PagerDuty

View all jobs at PagerDuty

Likely interview questions

  • Walk us through your experience operating Linux systems in production—what was the most complex incident you debugged?
  • Describe your hands-on experience with Kubernetes or EKS. What aspects of cluster operations are you most comfortable with?