Data Engineer - Onboarding
Job description
Data Engineer - Onboarding at sardine.
About the role
Sardine is seeking a Senior Data/ML Engineer to oversee the data and machine learning infrastructure that powers our compliance decisions. This role involves end-to-end ownership of data pipelines, from ingestion to feature creation, model development, and maintaining model accuracy in production. You will be a key individual contributor, shaping the technical direction for future growth in feature generation, KYC onboarding models, and internal entity matching for sanctions.
Key facts
What you'll do
- Manage the data ingestion layer for device telemetry, transaction events, KYC/identity signals, and third-party data, building correct, observable, and extensible streaming and batch pipelines.
- Develop and enhance our feature platform, ensuring consistent feature definitions are computed for both streaming and batch processes, with sub-second serving to the rules engine and models.
- Implement engineering practices to ensure feature accuracy, including streaming-versus-batch reconciliation, recomputation tests, and drift monitoring.
- Productionize fraud and identity ML models, creating training pipelines, using various gradient-boosted and tree-based models, and building automated retraining, promotion, and rollback systems.
- Engineer KYC, AML, and identity risk signals from diverse, multi-vendor, multi-jurisdiction data sources into usable model features.
- Integrate and fortify new data sources, including over 30 third-party enrichment providers and our cross-client consortium network, managing failover, timeouts, caching, and cost.
- Oversee the BigQuery warehouse and modeling layer, including partitioning, data staging, training datasets, and migrating from dbt to scheduled SQL and Python pipelines.
- Design entity resolution and graph data to link customers, devices, and financial accounts across clients, including large-scale connected-components analysis.
- Ensure platform security by design, implementing field-level encryption, regional data residency, PII handling, and feature-level gating.
- Provide technical leadership, writing design documents, conducting reviews, mentoring team members, and making build-versus-buy decisions.
Requirements
- 8 or more years of experience building production data and ML systems, with ownership of both pipeline and model development.
- Shipped models that have driven significant automated decisions, not just dashboards.
- Proficient in Python and SQL, with fluency in a distributed processing framework (Spark, Beam, or Flink) and understanding of streaming semantics.
- Hands-on experience with a modern cloud data stack, preferably GCP (BigQuery, Dataflow, Dataproc, Pub/Sub, Bigtable, Composer, Vertex AI) or AWS equivalents, plus Docker, Kubernetes, Terraform, and CI/CD.
- Practical ML engineering expertise, including feature stores, training/serving skew, gradient-boosted tree models, class imbalance, rare-event modeling, threshold tuning, model monitoring, drift detection, and explainability.
- Experience with high-volume, low-latency serving environments where feature fetches have tight latency budgets and no retry options.
- Domain experience in fraud, risk, payments, lending, or identity/KYC, or the ability to quickly become proficient in a regulated domain.
- Familiarity with data governance in a regulated environment, including PII, encryption, access control, regional data residency, and auditability.
- Strong written communication skills, capable of explaining technical concepts to both technical and non-technical audiences.
- Proactive approach and comfort with ambiguity, taking initiative to define and build solutions.
Nice to have
- Experience supporting customer-facing ML, such as bring-your-own-model integrations, model explainability for regulatory review, or shadow/challenger scoring frameworks.
- Experience in high-growth B2B SaaS, or as an early data/ML hire who established the function.
Skills & tools
- Python
- SQL
- Spark
- Apache Beam
- Flink
- GCP (BigQuery, Dataflow, Dataproc, Pub/Sub, Bigtable, Composer, Vertex AI)
- AWS (equivalents)
- Docker
- Kubernetes
- Terraform
- CI/CD
- XGBoost
- LightGBM
- CatBoost
- scikit-learn
- Airflow
- Chronon
- Kubeflow
Practical notes
This is a remote-first role for candidates in the United States or Canada. Benefits include health, dental, and vision insurance for employees and dependents, 4% matching in 401k/RRSP, a MacBook Pro, a home office stipend, and monthly stipends for meals, social meet-ups, health/wellness, and learning. We offer flexible paid time off and early exercise for all stock options.