Senior Data Engineer
Job description
About the role
You will own the design, build, and reliability of our cloud-native data lakehouse from raw ingestion through to analytics-ready Gold tables. You will work closely with data analysts, analytics engineers, and product stakeholders to deliver trusted data at speed while championing data quality and observability as first-class concerns. This role sits at the intersection of data engineering and platform engineering where you are expected to think in architectures not just pipelines. You will champion automated data validation and embed observability into every layer of the data platform. You will translate ambiguous business requirements into well-defined data models and pipeline designs in collaboration with analysts and stakeholders. You will mentor junior engineers on data engineering best practices and support the adoption of CI/CD for pipeline and dbt model deployment.
Key facts
What you'll do
Data Platform & Pipeline Engineering
▸ Design, build, and maintain scalable ETL/ELT pipelines using Azure Data Factory (ADF) and Apache Airflow, processing structured and semi-structured data across the Medallion architecture (Bronze → Silver → Gold).
▸ Implement incremental load patterns, change data capture (CDC), and event-driven ingestion to ensure data freshness across the platform.
▸ Build and optimise Snowflake data warehouse objects - tables, views, dynamic tables, streams, tasks, and stored procedures - for performance and cost efficiency.
▸ Develop modular, tested dbt models aligned to each Medallion layer, enforcing consistent naming conventions, documentation, and lineage across all transformations.
▸ Manage and optimise Azure Data Lake Storage Gen2 (ADLS) - folder structures, lifecycle policies, access tiers, and partition strategies.
▸ Build and maintain Azure Functions and Azure Logic Apps for lightweight event-driven processing, orchestration triggers, and operational automation.
▸ Manage secrets, credentials, and environment-specific configuration securely using Azure Key Vault - no hardcoded credentials in pipelines or code.
▸ Contribute to infrastructure-as-code practices for provisioning Azure data services using Terraform or Bicep preferred.
▸ Define and enforce data contracts between producers and consumers with row count checks, null rate thresholds, referential integrity, and value domain validation.
▸ Build and maintain data quality dashboards to give engineering and business stakeholders real-time confidence in platform health.
Requirements
Must-Have
▸ Snowflake: Advanced SQL - window functions, CTEs, recursive queries, query profiling.
▸ Snowflake: Native features - streams, tasks, snowpipe, dynamic tables, row-level security.
▸ Snowflake: Virtual warehouse tuning and credit cost optimisation.
▸ dbt + Elementary: Writing, testing, and documenting production dbt models.
▸ dbt + Elementary: Elementary integration for data observability and anomaly detection.
▸ dbt + Elementary: dbt incremental strategies, snapshots, and semantic layer.
▸ Azure Cloud: Azure Data Factory - pipeline authoring, triggers, parameterisation, linked services.
▸ Azure Cloud: ADLS Gen2 - zone/folder design, lifecycle management, Parquet/Delta partitioning.
▸ Azure Cloud: Azure Key Vault - secret management, managed identities.
▸ Azure Cloud: Azure Functions / Logic Apps - event-driven triggers and lightweight automation.
▸ Airflow: DAG authoring, task dependencies, XCom, sensors, and connection management.
▸ Airflow: Airflow deployment and monitoring in cloud-hosted environments.
▸ Python: Data pipeline scripting, PySpark basics, REST API integration.
▸ Python: Unit testing pipeline logic and transformation functions.
▸ Data Quality & Medallion Architecture: Hands-on experience implementing Bronze / Silver / Gold Medallion architecture.
▸ Data Quality & Medallion Architecture: Data validation checks at each layer - not just at the final Gold layer.
▸ Data Quality & Medallion Architecture: Schema evolution handling and SCD Type 2 dimension management.
▸ 4+ years of professional data engineering experience with at least 2 years on Azure cloud data platforms.
Nice-to-Have
▸ Exposure to Snowflake Cortex, dbt Semantic Layer, or Boomi Data Hub for AI-assisted data enrichment within pipeline layers.
▸ Experience integrating LLM-based quality checks or AI-assisted anomaly detection into data workflows.
▸ Familiarity with Microsoft Fabric and OneLake as a complementary or future-state platform.
▸ Knowledge of data mesh or data product thinking and how it maps to Medallion layer ownership.
▸ Experience with Terraform or Bicep for Azure infrastructure provisioning.
Practical notes
4+ years of professional data engineering experience required with at least 2 years on Azure cloud data platforms.