Staff Data Infrastructure Engineer
Job description
About the role
You define and own the core data infrastructure that powers decision-making and care delivery across Headway. You will architect the Snowflake data warehouse and design the ingestion and orchestration patterns that scale across all fifty states. You lead the evolution of our cloud infrastructure on AWS while ensuring robust monitoring, alerting, and security for sensitive mental health data. You set the technical standards for data modeling, CI/CD, and developer experience that the entire organization follows. You partner closely with product, analytics, and machine learning teams to turn complex requirements into reliable data products. You mentor engineers and elevate the craft through code reviews, documentation, and thoughtful system design. You are the technical anchor who enables non-technical staff to safely leverage AI-driven workflows built on a solid data foundation.
Key facts
What you'll do
Architect and evolve the foundational data warehouse, ingestion layer, and orchestration systems on Snowflake and AWS.
Define and enforce data platform standards, schemas, and API contracts to ensure consistency across teams and products.
Lead the design of scalable ingestion pipelines that reliably integrate with a wide network of providers, payers, and internal systems.
Implement robust monitoring, alerting, and logging frameworks to maintain high reliability and fast incident response.
Build and maintain CI/CD workflows and infrastructure-as-code tooling to enable safe, repeatable deployments at high velocity.
Optimize query performance and cost efficiency for complex analytical workloads on a large MPP database.
Partner with data analysts and machine learning teams to deliver curated datasets and features that accelerate their work.
Design and operate data security, compliance, and access control strategies to protect sensitive mental health information.
Mentor data engineers and collaborate cross-functionally to codify best practices and elevate the overall engineering bar.
Evaluate and integrate third-party data tools and services to improve platform capabilities and reduce operational burden.
Champion observability and data quality frameworks to ensure trust, transparency, and reliability in all data outputs.
Lead end-to-end technical initiatives from discovery through production rollout, coordinating with stakeholders across the organization.
Drive automation of data infrastructure tasks to reduce manual effort and increase platform self-service.
Requirements
10+ years as a Data Platform Engineer, Software/Infrastructure Engineer specializing in data, or Data Engineer in a high-code, high-scale environment.
Track record of driving technical strategy and cross-functional alignment, including defining roadmaps for a data platform or infrastructure domain.
Deep expertise building full-stack data platforms: data warehousing, ingestion pipelines, orchestration, monitoring and alerting, CI/CD, developer tooling, cloud infrastructure, and third-party integrations.
Architectural fluency in warehouse design patterns, performance tuning, cost management, and permissions strategies on MPP analytics databases such as Snowflake, with Databricks, BigQuery, or Redshift also relevant.
Experience leading technical initiatives end-to-end across multiple teams and codifying engineering standards that persist beyond individual projects.
Proficiency designing and operating cloud infrastructure at scale on AWS using infrastructure-as-code tools such as Terraform, Pulumi, or AWS CDK.
Strong Python engineering skills combined with solid SQL foundations and comfort with distributed systems concepts and trade-offs.
Experience maintaining and scaling pipeline orchestration infrastructure, with a preference for tools like Airflow or Astronomer.
Demonstrated mentorship track record growing junior and mid-level engineers through guidance, feedback, and technical leadership.
Experience with data security and compliance in regulated environments, particularly around protected health information, is essential.
Nice to have: Agentic Data Engineering practices such as automating data infrastructure tasks, auditing permissions, and optimizing CI workflows.
Nice to have: Proficiency with APM and observability tooling such as DataDog or New Relic, and data quality and observability frameworks.
Nice to have: Deep experience with ETL/ELT best practices at scale and proven ability to make build vs. buy decisions and evaluate vendors.
Nice to have: Background in greenfield platform development within a high-growth startup environment.
Nice to have: Hands-on experience with Docker, GitHub Actions, dbt, and Spark for data processing at scale.
Practical notes
This role is based in New York, NY and is a full-time engagement.
The expected base pay range for this position is $212,000 - $265,000, based on a variety of factors including qualifications, experience, and geographic location. In addition to base salary, this role may be eligible for an equity grant, depending on the position and level.
We are committed to offering an accessible and inclusive work environment where all team members can bring their whole selves to work.