Twilio
Software Engineer-Platform Engineering (L3)
About this role
Twilio seeks an L3 Software Engineer for its Platform Engineering team to design, build, and operate highly available distributed systems handling billions of emails at scale. This remote role requires deep expertise in Kubernetes, cloud infrastructure across AWS and Azure, and a commitment to code quality and system reliability.
What you'll do
- Design and operate Kubernetes clusters at scale, partnering with product and leadership on complex system requirements
- Drive code reviews and enforce maintainable patterns with rigorous testing standards across unit, integration, and component levels
- Manage cloud infrastructure across AWS and Azure using Terraform, ensuring observability through metrics, alerts, and distributed tracing
- Identify and mitigate technical debt, system bottlenecks, and single points of failure while balancing feature delivery
- Mentor junior engineers and lead technical sprint planning across distributed teams
- Operate resilient backend services under production load using modern AI-assisted development tools
What they're looking for
- Kubernetes cluster management and zero-downtime upgrades
- AWS and Azure cloud infrastructure
- Terraform and Infrastructure as Code
- GitOps and ArgoCD
- Distributed systems design and operation
- Backend service development
- AI-assisted development tools (e.g., Claude Code)
- System observability and monitoring
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Twilio
Twilio builds communication and identity infrastructure platforms including voice, email, and authentication services, enabling enterprises to integrate reliable messaging and identity solutions at scale. The company is hiring software engineers, forward-deployed engineers, and senior engineers to develop scalable backend systems, design cloud infrastructure, implement AI-powered features, and partner with customers on production deployments.
- Website
- twilio.com
Likely interview questions
- Describe your experience managing Kubernetes cluster upgrades with zero downtime—how have you handled node draining and PodDisruptionBudgets?
- Tell us about a time you identified and resolved a critical single point of failure in a large-scale distributed system.