Astronomer
Customer Reliability Engineer, Infrastructure
About this role
Join Astronomer's Customer Reliability Engineering team as an Infrastructure Specialist focused on maintaining and monitoring cloud infrastructure and Kubernetes clusters for our managed Airflow service. You'll be directly responsible for customer success, incident response, and building observability systems while working with leading enterprises across multiple cloud providers.
What you'll do
- Respond to and resolve infrastructure incidents from customers and monitoring systems
- Troubleshoot and triage customer environments while maintaining SLAs
- Build and maintain monitoring, alerting, and observability systems
- Develop automation for operational tasks and daily maintenance
- Provide direct customer engagement and guidance on production deployment
- Participate in on-call rotation including weekend coverage
What they're looking for
- Kubernetes (3+ years)
- Cloud infrastructure (AWS, GCP, or Azure)
- Linux administration
- Python scripting
- Distributed systems monitoring and troubleshooting
- DevOps and CI/CD practices
- Customer support and communication
- Infrastructure as Code (bonus)
Benefits
- Remote work opportunity across the United States
- Equity component
- Comprehensive benefits package
- Exposure to diverse industries and multi-cloud environments
- Direct impact on customer success and product improvements
- Opportunity to work with cutting-edge data orchestration technology
Opens the application — the Jobs AI extension fills it for you. Set up autofill
Opens the official application on the employer’s site. No login required.
Astronomer
Astronomer builds a DataOps platform powered by Apache Airflow that helps enterprises manage data workflows at scale. The company is hiring for technical customer-facing roles including Customer Reliability Engineers, Sales Engineers, Field Engineers, and Infrastructure Specialists who support enterprise customers' Airflow adoption and platform reliability.
- Website
- astronomer.io
Likely interview questions
- Describe your experience managing production Kubernetes clusters at scale and how you've handled critical incidents.
- Walk us through a complex troubleshooting scenario you've encountered in a distributed system and how you resolved it.