Customer Reliability Engineer
Job description
About the role
This role operates within the Astronomer Customer Reliability Engineering team to ensure success and reliability for customers using the managed Airflow service. The position focuses on cloud and Kubernetes infrastructure reliability while responding to incidents and supporting customers directly.
Platform and DevOps engineers build the systems that run everything else. They manage infrastructure, CI/CD pipelines, observability, and reliability. The work is about automation, scaling, and removing friction for product teams. Systems thinking is the core skill. Platform teams are measured by developer velocity and system reliability. Most companies run on-call rotations, and understanding incident response is part of the role.
Key facts
What you'll do
Operate, monitor, and maintain the managed platform to ensure availability and predictable performance for customers using Astronomer products.
Provide solutions that help customers succeed with Astronomer products and meet their expectations across industries, including radio and audio campaigns.
Triage customer environments by investigating issues and coordinating active resolutions with clients.
Own the observability platform and drive permanent fixes or monitoring for problems identified through monitoring systems.
Handle incidents raised by customers or detection systems and implement lasting resolutions or safeguards.
Support weekend coverage through on-call rotation for critical platform events.
Supply input to product teams by articulating customer needs, pain points, and emerging requirements.
Construct and refine monitoring and alerting systems to improve response and detection accuracy for observability platforms.
Automate recurring operational tasks to increase efficiency and reduce manual effort for daily operational tasks.
Guide product architecture decisions and contribute where appropriate based on operational insights into platforms such as SDK, Android, and firmware builds.
Champion the customer experience by prioritizing issues, meeting SLAs, and guiding clients toward production readiness.
Participate fully in a distributed, remote team using remote collaboration practices.
Enhance customer documentation to improve clarity and self-service support for products.
Work with modern technology stacks and multi-cloud provider implementations, including AWS, GCP, and Azure.
Requirements
Holders must bring 5 years of experience with large, complex cloud infrastructures operating at scale.
Holders must include 3 years of hands-on Kubernetes experience.
Holders must manage production distributed systems on at least one major cloud provider such as AWS, GCP, or Azure.
Holders must demonstrate strong Linux proficiency in production contexts.
Holders must understand how to operate and monitor distributed systems in production.
Holders must have prior experience handling customer issues in internal or external settings.
Holders must have DevOps or CI/CD experience to support platform workflows.
Holders must be able to write Python scripts for automation and tooling.
Holders must show strong troubleshooting skills across distributed environments.
Practical notes
Success in this role requires comfort with remote work in a fully distributed team.
The position may involve travel or remote participation based on team and customer needs.
Typical interview steps
Platform interviews usually include an infrastructure scenario, a scripting or coding exercise, and operational questions. Candidates may be asked to design a deployment pipeline or debug an outage. Incident experience and an automation mindset are tested. Interviewers often ask about a past outage and how you handled it. Structured post-incident thinking, not heroics, is what they look for.
Good to know
Reliability engineering focuses on ensuring systems remain available and performant for users.
Site reliability practices help balance software release velocity with stability and risk management.
Infrastructure as code tools are commonly used to manage cloud environments reproducibly.
Observability platforms combine metrics, logs, and traces to surface meaningful signals.
Apache Airflow is an open source workflow orchestration platform often used for data pipelines.
About the company
Astronomer empowers data teams to bring mission-critical software, analytics, and AI to life and is the company behind Astro, the industry-leading unified DataOps platform powered by Apache Airflow®. Astro accelerates building reliable data products that unlock insights, unleash AI value, and powers data-driven applications.