
Senior Data Engineer
Job description
About the role
At Ambience, we are constructing the AI intelligence platform that restores humanity to healthcare and drives meaningful ROI for health systems across the country. The role centers on solving a hard data engineering problem that directly impacts how quickly and accurately the company can deliver verified answers to critical business questions. You will own the pipelines that ingest clinical and operational data, model it into governed metrics, and deliver trusted self-service analytics. The hire will work cross-functionally with product managers, engineers, clinicians, and GTM teams to turn raw data into the numbers the entire organization relies on for decision-making. If a metric is wrong, stale, or cannot be explained, this role owns the fix, and the credibility of the platform depends on the execution. This is a hybrid position requiring presence in the San Francisco office three days per week.
Key facts
What you'll do
- Own the Trust Layer by building and running the pipelines that ingest, validate, and transform clinical and operational data at scale to ensure metrics are accurate, reproducible, and defensible.
- Enable Insights Across Teams by partnering with engineering, clinical, and product stakeholders to design clear, actionable dashboards and analytics that guide decision-making and improve healthcare outcomes.
- Support Scalable Infrastructure by applying best practices around warehousing, orchestration (Dagster), governance, and RBAC to keep data systems secure, performant, and ready for rapid innovation as load increases.
- Automate Data Ingestion Workflows by building file-based ingestion pipelines that enable plug-and-play onboarding of external data sources and developing automated validation, triggering, and error-handling mechanisms for real-time data availability.
- Establish Validation & Transformation Pipelines by implementing data validation frameworks using Python or TypeScript, and designing transformation layers (SQL, dbt) that standardize and cleanse raw data for analysis and operational workflows.
- Deliver Governed, Trusted Analytics by building and maintaining governed metrics and analytical schemas in SQL that stand up to customer and third-party verification.
- Enable tiered self-service so every employee - technical or not - can answer their own questions without waiting on the data team, reducing bottlenecks and accelerating insights.
- Collaborate cross-functionally to translate ambiguous business problems into data models and pipelines, ensuring alignment between data outputs and strategic objectives.
- Maintain end-to-end ownership of production pipelines, including observability, monitoring, and rapid response to issues, understanding that data downtime directly impacts customer trust and internal decision quality.
- Drive continuous improvement by identifying inefficiencies in current workflows and proposing scalable solutions that balance speed with reliability and governance.
Requirements
- 5+ years in a production Data Engineering role with hands-on Snowflake experience, including data modeling, performance tuning, and warehouse administration.
- Strong SQL skills and proficiency in Python to build robust, maintainable data pipelines.
- Experience building ETL/ELT pipelines, data lakes, or warehouses in modern cloud environments.
- Solid grasp of data validation techniques, schema design, and scalable architecture to support growing data volumes and complexity.
- A clear communicator who can bridge technical and non-technical teams, gathering requirements and presenting insights to cross-functional stakeholders.
- Mission-driven mindset that thrives in a fast-paced startup environment and takes full ownership of deliverables and outcomes.
- Comfort with making decisions and driving execution without needing a fully-specified ticket, understanding that speed is critical in a growth-stage company.
- Proven ability to own a production pipeline end-to-end and recognize the importance of observability because you have been paged when data went stale.
- Someone who moves fast, takes real ownership of data quality, and is comfortable calling the shots when specifications are incomplete.
Nice to have
Early-stage startup experience, background in regulated industries such as healthcare or finance, and familiarity with dbt, orchestration tools like Airflow or Dagster, or BI tools such as Looker, Tableau, or Mode.
Practical notes
This role is hybrid in the San Francisco office three days per week.