Principal Software Engineer
Job description
About the role
You will own the technical strategy and execution for a cloud-agnostic, self-contained data platform that allows enterprises to run Apache Airflow entirely within their own infrastructure. This role requires you to architect and deliver both the control plane and data plane as completely isolated, on-premises solutions that operate without any dependency on a SaaS backbone or mandatory outbound communication. You will define the end-to-end architecture that ensures reliability and performance across diverse network topologies, from standard cloud deployments to strictly air-gapped environments. You will serve as the primary technical authority and public-facing architect, translating complex enterprise requirements into coherent engineering roadmaps. Additionally, you will mentor senior engineers, drive design excellence, and ensure that the platform enables customers to unlock insights and power data-driven applications reliably and securely.
Key facts
What you'll do
- Define, Design, and Build Core Architecture: Own the end-to-end definition, design, and implementation of both the control plane and data plane services, creating highly resilient inter-connections between them to power this self-contained infrastructure platform.
- Architect an Extensible Platform API: Drive the design, data modeling, versioning, and long-term backward compatibility of APIs (such as REST or gRPC). These APIs serve as the core platform engine that power the customer-facing UI, and the developer endpoints for customer-facing programmatic integrations.
- Design Comprehensive Platform Observability: Architect the telemetry framework (metrics, logs, and traces) for the self-contained platform itself, ensuring deep visibility into both the core control plane services and the performance, health, and status of the underlying Airflow instances it provisions.
- Architect for Varied Isolation Levels: Oversee the architecture of a self-contained software platform, ensuring all services, images, dependencies, and configurations can run smoothly whether deployed in connected private clouds or highly secure, zero-egress networks.
- Provide Technical Leadership: Act as the technical face of the team, translating complex enterprise infrastructure needs into execution roadmaps while aligning closely with product management and executive stakeholders.
- Mentor and Elevate: Foster a culture of engineering and operational excellence through comprehensive design reviews, high- code contributions, and direct career mentorship for senior engineers.
- Define and Enforce Platform Security Boundaries: Establish and maintain the security models and controls necessary for a self-contained system that operates in customer-controlled network environments without external dependencies.
- Guide Technology Adoption: Evaluate, prototype, and standardize tools and frameworks across the stack to ensure the platform remains performant, maintainable, and aligned with enterprise deployment constraints.
Requirements
- 10+ years of professional software engineering experience with a significant track record operating at a Staff, Senior Staff, or Principal level within an enterprise infrastructure, database-as-a-service (DBaaS), or private cloud organization.
- Proven success building production-grade infrastructure control planes, with a deep understanding of API design principles required to support programmatic customer workflows and user interfaces.
- Practical experience designing and implementing infrastructure observability stacks (e.g., Prometheus, OpenTelemetry, Grafana) to capture, aggregate, and surface system health and workflow data within isolated environments.
- Solid operational knowledge and architectural understanding of container orchestration platforms, specifically Kubernetes. You have practical experience building custom Kubernetes Operators (CRDs), managing Helm deployments, and working with the Kubernetes API to orchestrate application lifecycles.
- Hands-on systems design expertise involving isolated environments across multiple public cloud providers (AWS, GCP, Azure), Red Hat OpenShift enterprise deployments, and bare-metal on-premise infrastructure.
- Expert-level software development skills (e.g., Go, Python, TypeScript) with deep knowledge of distributed state stores (such as Etcd, Consul, or PostgreSQL), concurrency, and network security patterns.
- Exceptional oral and written communication skills with a proven ability to author crisp, high-density technical "one-pagers" and actionable RFCs. You can drive fast organizational alignment by distilling complex systems choices into brief, highly readable documents.
- Demonstrated ability to operate in regulated and strict security environments, designing systems that meet enterprise compliance and audit requirements without reliance on external services.