Senior Data Engineer (GCP
Job description
About the role
You will be a foundational, greenfield hire owning the complete design and implementation of our analytical data platform on Google Cloud Platform from day one. This role demands a hands-on builder who transforms raw infrastructure into a trusted, governed data core that powers both internal decision intelligence and external customer features. You will own the data and retrieval layer end-to-end, enabling our internal Claude-based agentic systems to reason reliably on curated company information. In parallel, you will construct modeled data pipelines and feature stores that directly power predictive and prescriptive insights visible on customer dashboards. You will work in close partnership with our Senior AI Engineer, ensuring the data and retrieval substrates align with agentic logic and orchestration requirements. This position requires a pragmatic expert who balances robust governance with rapid, production-grade delivery in a collaborative environment. You will define standards for metrics, dimensional modeling, and semantic clarity so that data serves both humans and machines with consistency.
Key facts
What you'll do
Design and build the first enterprise-grade analytical data platform on GCP, including warehouse, ingestion layer, transformation framework, and orchestration tooling.
Construct robust ELT and CDC data pipelines sourcing from our production Postgres database and .NET/C# services using GCP-native services such as Datastream, Dataflow, and Pub/Sub.
Establish end-to-end orchestration workflows using Cloud Composer or Airflow, and implement data CI/CD environments with automated testing and deployment pipelines.
Create dimensional models and a semantic layer with clear metric definitions to serve analytics dashboards, data analysts, and internal AI agents.
Build and maintain the retrieval infrastructure for embeddings, vector stores, and RAG plumbing to keep internal and external agents grounded in timely, accurate data.
Own data quality, observability, lineage, and PII governance controls across sensitive financial, donor, payments, and cross-border tax data sets.
Implement scalable streaming and change data capture patterns to support near-real-time data serving and agent use cases.
Collaborate with the AI Systems Engineer and .NET backend team to instrument product events and ship customer-facing insights through reliable data products.
Leverage expert SQL and PostgreSQL skills alongside cloud data warehouse technologies, primarily BigQuery, to ensure performant and maintainable solutions.
Write production-grade Python code to develop custom pipeline components, tooling, and integrations with the GCP data stack.
Demonstrate experience with modern LLM data patterns including embeddings pipelines, vector databases such as pgvector or Vertex AI Vector Search, and feature stores.
Champion data governance, security, and compliance practices aligned with regulatory requirements and internal risk policies.
Requirements
Bring a minimum of 6 years of professional data engineering experience, with at least one greenfield warehouse or platform build that you owned from design through production.
Possess expert-level SQL skills and deep hands-on experience with PostgreSQL, along with strong proficiency in cloud data warehouses, BigQuery being a clear ideal.
Show advanced capability in ELT and ETL design using dbt combined with a modern orchestrator such as Airflow, Cloud Composer, or Dagster.
Have concrete experience with the GCP data stack, including BigQuery, Datastream, Dataflow, Pub/Sub, and Cloud Storage for storage and processing.
Be proficient in Python for data pipeline development, custom tooling, and integration work within cloud environments.
Demonstrate proven experience building and maintaining retrieval substrates for LLMs, including embeddings pipelines, vector stores like pgvector or Vertex AI Vector Search, RAG plumbing, and feature stores serving production ML or LLM systems.
Show a strong understanding of data quality frameworks, data lineage, observability, and PII governance, especially around sensitive financial, donor, payments, and cross-jurisdictional tax data.
Have a track record of collaborating with software engineers, AI specialists, and product teams to deliver reliable data products under production constraints.
Nice to have
Familiarity ingesting data from .NET/C# backends and tracking events from web and mobile clients built with React or React Native.
Experience with streaming architectures, change data capture at scale, and near-real-time data serving patterns.
Background in ML feature engineering, MLOps, and the Vertex AI ecosystem for model and feature lifecycle management.
A history of working in fintech, payments, or the charitable-giving and nonprofit sectors.
Direct experience constructing data layers that production LLM agents rely on for timely and accurate responses.
Practical notes
This is a full-time position open to remote or hybrid work across all cities.
Employees may be eligible for additional holidays after completing 1 year and again after 5 years of service.
The role supports sponsored training and certification programs to support continuous professional development.
An employee referral bonus program is available to eligible staff.
Comprehensive private health insurance is provided, including dental care.
A multisport card is offered and fully covered by the company.
The office environment is designed to be fun and collaborative, with relaxation zones and free parking available on-site.