Skip to main content

Astronomer

Customer Reliability Engineer, Infrastructure

Remote (United States) (Remote)$125k–$130kfulltimemidAdded today

About this role

Join Astronomer's Customer Reliability Engineering team as an Infrastructure Specialist focused on maintaining and monitoring cloud infrastructure and Kubernetes clusters for our managed Airflow service. You'll be directly responsible for customer success, incident response, and building observability systems while working with leading enterprises across multiple cloud providers.

What you'll do

  • Respond to and resolve infrastructure incidents from customers and monitoring systems
  • Troubleshoot and triage customer environments while maintaining SLAs
  • Build and maintain monitoring, alerting, and observability systems
  • Develop automation for operational tasks and daily maintenance
  • Provide direct customer engagement and guidance on production deployment
  • Participate in on-call rotation including weekend coverage

What they're looking for

  • Kubernetes (3+ years)
  • Cloud infrastructure (AWS, GCP, or Azure)
  • Linux administration
  • Python scripting
  • Distributed systems monitoring and troubleshooting
  • DevOps and CI/CD practices
  • Customer support and communication
  • Infrastructure as Code (bonus)

Benefits

  • Remote work opportunity across the United States
  • Equity component
  • Comprehensive benefits package
  • Exposure to diverse industries and multi-cloud environments
  • Direct impact on customer success and product improvements
  • Opportunity to work with cutting-edge data orchestration technology
Apply with Autofill

Opens the application — the Jobs AI extension fills it for you. Set up autofill

Opens the official application on the employer’s site. No login required.

Astronomer

Astronomer builds a DataOps platform powered by Apache Airflow that helps enterprises manage data workflows at scale. The company is hiring for technical customer-facing roles including Customer Reliability Engineers, Sales Engineers, Field Engineers, and Infrastructure Specialists who support enterprise customers' Airflow adoption and platform reliability.

View all jobs at Astronomer

Likely interview questions

  • Describe your experience managing production Kubernetes clusters at scale and how you've handled critical incidents.
  • Walk us through a complex troubleshooting scenario you've encountered in a distributed system and how you resolved it.