Kong
Site Reliability Engineer 2
About this role
Kong is seeking a Site Reliability Engineer to build and operate their multi-region SaaS platform (Konnect) across AWS, GCP, and Azure. You'll manage Kubernetes clusters, service mesh architectures, and complex data infrastructure while ensuring reliability and security at scale for thousands of customers worldwide.
What you'll do
- Operate and scale Kong's global SaaS platform ensuring reliability, availability, and performance across multiple regions and cloud providers
- Build and maintain Kubernetes-based infrastructure using Terraform, Helm, and ArgoCD with GitOps workflows
- Design and optimize multi-region data and caching layers including PostgreSQL, Redis, ClickHouse, and Druid
- Develop CI/CD pipelines and automate service delivery while maintaining consistent infrastructure changes
- Enhance observability and incident response using Datadog, Prometheus, Grafana, and Thanos; define and track SLOs
- Participate in 24/7 on-call rotation and drive continuous improvement through scaling initiatives and postmortem practices
What they're looking for
- Kubernetes administration and debugging
- Infrastructure as Code (Terraform/Terragrunt)
- CI/CD and GitOps (ArgoCD, Helm)
- Programming languages (Go, Python, Bash)
- Linux/Unix systems and networking (DNS, TLS/SSL, HTTP)
- Observability platforms (Datadog, Prometheus, Grafana)
- API gateway and service mesh technologies
- Distributed systems and load balancing
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Kong
Kong builds an API and AI gateway platform for cloud-native environments that enables enterprise customers to manage connectivity and integrations. The company is hiring Forward Deployed Engineers who combine hands-on software engineering with direct customer collaboration to implement, optimize, and migrate to Kong's solutions.
- Website
- konghq.com
Likely interview questions
- Walk us through your experience managing SaaS or PaaS systems at enterprise scale—what was the largest production incident you handled and how did you resolve it?
- Describe a complex multi-region Kubernetes deployment you designed—what were the key challenges and how did you ensure fault tolerance?