Staff Software Engineer, Observability
Job description
About the role
At Astronomer, our research and development organization is dedicated to providing an exceptional experience in operating data orchestration based on Apache Airflow at many of the world's largest companies. As we continue to grow our platform, we are seeking to solve highly complex technological challenges through inventive algorithms and best practices around scale. The company is building out a Observability team to deliver data observability capabilities that give our customers visibility, reliability, and actionable insights into their data pipelines and products across some of the largest enterprises globally. In this role, you will contribute directly to the design, development, and scaling of Astronomer's observability platform, working closely with product, design, and cross-functional engineering teams. Your work will have a significant impact on our customers' experience and the future development of the platform, helping to ensure data reliability and operational excellence.
Key facts
Location: USA
Engagement: Full-time
The estimated salary for this role ranges from $275,000 to $377,000, depending on leveling and geographic factors. The compensation package includes salary, equity, and a comprehensive benefits package. This role is on-site in New York City, and the company values diversity and is an equal opportunity employer.
What you'll do
- Lead the end-to-end architecture and evolution of major components within the observability platform, making foundational design decisions that will influence the platform's long-term direction. This involves designing scalable, reliable, and high-performance features that provide deep visibility into customer data pipelines and operational metrics.
- Build and enhance features that enable customers to monitor, troubleshoot, and optimize their data workflows effectively, ensuring the platform can handle large-scale data and high throughput environments.
- Collaborate closely with product managers, designers, and engineering teams to define project scope, prioritize initiatives, and deliver high-impact solutions aligned with customer needs and business goals.
- Write high-quality, maintainable code following best practices in testing, continuous integration/continuous deployment (CI/CD), and operational excellence to ensure platform stability and reliability.
- Establish and promote engineering standards related to code quality, testing, system design, and operational procedures across the team, fostering a culture of excellence and continuous improvement.
- Improve and evolve tooling, infrastructure, and processes that support the observability platform, including monitoring, alerting, and incident response systems.
- Lead and coordinate responses to complex production incidents, guiding the team through rapid resolution while identifying root causes and implementing long-term fixes to prevent recurrence.
- Proactively identify technical opportunities, risks, and gaps within the platform, and drive initiatives to address these areas to improve stability, scalability, and usability.
- Mentor engineers across all levels through code reviews, pairing sessions, and providing constructive feedback, supporting their professional growth and fostering a collaborative engineering culture.
- Contribute to the development of best practices and documentation to ensure knowledge sharing and onboarding for new team members.
- Engage with customers and internal stakeholders to understand their observability needs and translate those into technical requirements and solutions.
- Stay current with industry trends and emerging technologies related to data observability, distributed systems, and cloud infrastructure to continuously enhance the platform's capabilities.
Requirements
- Extensive experience designing, developing, and maintaining complex distributed systems, with a strong preference for experience with Golang, Python, Kubernetes, SQL/OLAP databases, Kafka, and stream processing technologies.
- Proven track record of building scalable, reliable production infrastructure and features that support high data throughput and operational resilience.
- Deep understanding of data pipelines, observability, monitoring, and infrastructure at scale, with familiarity with tools such as Apache Airflow or similar workflow orchestrators.
- Strong collaboration and communication skills, with the ability to work effectively across cross-functional teams including product, design, and engineering.
- Demonstrated ability to deliver customer-focused solutions by translating technical requirements into impactful features and improvements.
- Enthusiasm for fostering an inclusive, healthy engineering culture, supporting team members' growth, and encouraging best practices.
- Ability to work effectively in ambiguous environments, breaking down complex open-ended problems into structured, actionable plans.
- Experience leading incident response efforts, troubleshooting production issues, and implementing long-term solutions.
- Familiarity with CI/CD pipelines, testing frameworks, and operational tooling to ensure high-quality releases and platform stability.
- Strong problem-solving skills with a focus on reliability, scalability, and performance.
Nice to have
- Experience building or scaling data observability products, with knowledge of industry leaders such as Datadog, Honeycomb, or similar platforms.
- Background working with Apache Airflow, or similar workflow orchestration tools, and understanding of their integration with observability solutions.
- Prior experience in monitoring, observability, or infrastructure engineering at large-scale organizations.
- Direct experience working with product teams to develop customer-facing features and solutions.
- Knowledge of AI and machine learning applications in the context of data observability and platform automation.
Skills & tools
- Python, Golang, Kubernetes, SQL, Kafka, CI/CD, AI, Apache Airflow, R
Practical notes
- This is an on-site role based in New York City, requiring physical presence at the office.
- The role involves working with complex distributed systems, large-scale data pipelines, and infrastructure.
- The compensation package includes salary, equity, and benefits, aligned with industry standards.
- The company values diversity and is committed to creating an inclusive environment for all employees.
- Candidates should be prepared to engage in collaborative problem-solving, incident response, and ongoing platform improvements.
- The role offers opportunities to influence the future of data observability and work with technologies in the data ecosystem.
About the company
Astronomer empowers data teams to bring mission-critical software, analytics, and AI to life and is the company behind Astro, the industry-leading unified DataOps platform powered by Apache Airflow®. Astro accelerates building reliable data products that unlock insights, unleash AI value, and powers data-driven applications.