Staff Software Engineer, Platform Infrastructure
Job description
About the role
You will own the design and delivery of the foundational platform systems that underpin Astronomer's products Astro, Observe, and our IDE. You will translate platform strategy into production-grade code while reasoning deeply about failure modes, latency budgets, and scalability constraints. This role requires you to make high-impact architectural decisions and be directly accountable for their outcomes across a multi-cloud infrastructure estate. You will partner with product and infrastructure teams to define requirements and then implement the systems that enable hundreds of enterprise customers to deliver data workloads reliably. Your work will be visible and consequential, shaping the capabilities and reliability of Astronomer's platform for years to come. You will be expected to operate systems in production, respond to critical incidents, and mentor engineers on best practices for building robust platform services.
Key facts
What you'll do
Define and drive the platform infrastructure strategy across multiple cloud providers, taking end-to-end ownership from discovery through implementation and operations.
Evaluate build versus buy tradeoffs for storage, networking, and orchestration components, and make reasoned recommendations aligned with reliability, cost, and operational complexity.
Design and implement control plane services that manage Kubernetes clusters, scheduling decisions, and multi-tenant resource allocation at scale.
Create and maintain comprehensive architecture decision records, ensuring that the rationale for key platform choices is transparent and accessible to the organization.
Collaborate closely with product teams to translate platform constraints and opportunities into clear requirements that enable efficient product development.
Lead incident postmortems for platform outages, extracting learnings and driving changes that reduce future failure risks and improve system resilience.
Partner with security and compliance stakeholders to ensure platform components meet enterprise standards for access control, auditability, and data protection.
Mentor engineers across the platform team by providing code reviews, design guidance, and hands-on contributions to critical services written in Go.
Operate and tune production Kubernetes clusters, diagnosing performance bottlenecks and applying fixes that improve stability and utilization across cloud environments.
Champion observability practices by instrumenting platform services, defining meaningful SLIs and SLOs, and driving improvements against reliability targets.
Evaluate and integrate storage primitives at the system level, choosing between relational, object, and specialized stores based on workload characteristics and failure modes.
Work cross-functionally with distributed systems experts to ensure platform components scale predictably under growing customer demand and workload complexity.
Represent Astronomer in technical forums and external discussions, helping to shape the direction of the DataOps ecosystem and the Apache Airflow integration roadmap.
Lead the design of multi-cloud networking and connectivity solutions that balance performance, security, and cost across AWS, GCP, and Azure.
Build and maintain internal tools that enable platform and product teams to deploy, monitor, and debug workloads with confidence and minimal friction.
Requirements
You have built production platform systems at scale and understand what it means to be responsible for their reliability and performance.
You possess distributed systems depth, including consistency and availability tradeoffs, failure cascades, backpressure, and graceful degradation patterns.
You apply non-abstract design thinking to new systems, creating diagrams, evaluating failure modes, and prioritizing mitigations based on real impact.
You have strong operational experience with Kubernetes, including control plane behavior and scheduler decisions under load.
You are fluent in Go and have a track record of shipping production services that are well-structured, testable, and performant.
You have multi-cloud experience across AWS, GCP, and/or Azure, having evaluated and lived with the consequences of architectural decisions in production.
You can define technical requirements and influence technology choices across a large engineering organization.
You communicate effectively in writing and speaking, producing design documents that change thinking and postmortems that drive improvement.
You have worked effectively in globally distributed teams, aligning with colleagues across time zones and cultures.
Nice to have
Experience with storage primitives at the system level, including reasoning about when to use relational, object, or other stores.
Experience building or operating a SaaS or PaaS product across multiple cloud providers.
Practical notes
This role is full-time based in New York City.
Candidates must be eligible to work in the United States without sponsorship.
No relocation or visa sponsorship is available for this position.
The work location and compensation details are subject to internal policies and applicable laws.