Senior Data Engineer
Job description
About the role
You architect and own the core ingestion and normalization pipelines that feed the Shadow AI coordination layer on a daily basis. You translate messy, real-world marketing data into reliable, queryable structures that hundreds of concurrent users depend on. You are responsible for ensuring data moves securely and accurately from third-party sources to our warehouse with minimal manual intervention. You partner closely with AI and data science teams to guarantee the data they retrieve is consistent, well-modeled, and trustworthy. You influence engineering standards and tooling choices that shape how the entire data team operates. You engage directly with enterprise customers to understand their data constraints and surface opportunities for improvement. You treat data reliability as a product feature, not an afterthought.
Key facts
What you'll do
Design and maintain scalable ingestion workflows for marketing APIs including Meta, Google, TikTok, GA4, Shopify, and Klaviyo with robust auth and rate-limit handling.
Implement incremental sync strategies and backfill mechanisms that keep large-scale, multi-tenant datasets accurate and up to date.
Define and enforce normalized schemas that unify campaign, creative, and order data across disparate marketing platforms into a cohesive model.
Build observability and data quality systems that detect sync failures, anomalies, and drifts before downstream users are impacted.
Optimize pipeline performance and cost efficiency to support thousands of connections and hundreds of brands without degradation.
Collaborate with AI and data science teams to expose clean, retrieval-friendly datasets that power agent memory and reasoning.
Own the security and compliance posture of data pipelines, ensuring strict isolation, encryption, and auditability for enterprise customers.
Lead operational runbooks and incident response for data infrastructure in production environments with high availability requirements.
Champion best practices in schema design, versioning, and documentation to enable long-term maintainability and team scalability.
Mentor and influence engineering culture by setting standards for testing, monitoring, and reliability in data-intensive systems.
Requirements
You have built and operated large-scale data pipelines that served 1,000+ users or handled equivalent data volume where reliability and isolation were critical constraints.
You possess strong SQL and Python skills and have shipped production workloads in a modern data warehouse such as BigQuery, Snowflake, or Redshift.
You deeply understand ETL and ELT patterns, incremental synchronization, schema evolution, and dimensional modeling for analytics workloads.
You have hands-on experience integrating with third-party APIs, managing OAuth flows, handling pagination, and adapting to schema drift in production.
You are obsessive about observability, instrumentation, and data quality, and you refuse to ship data unless you can trust its correctness and lineage.
You have experience designing and operating systems that meet SOC 2 security standards, including access controls, encryption, data isolation, and auditability.
You are comfortable working in a fully remote setting with overlapping hours in the EST time zone (9 AM to 5 PM).
You are based in Pakistan and able to work full-time under an employment engagement without relocation.
Nice to have
You have worked in martech, adtech, or other data-heavy marketing domains and understand the nuances of attribution windows, timezone handling, and deduplication across platforms.
You are familiar with our core stack, including GCP (BigQuery, Cloud Run), PostgreSQL with pgvector, and orchestration tools such as dbt, Airflow, or Dagster.
You have built pipeline observability and tracing in AI or LLM contexts, for example using Langfuse or similar platforms.
You have experience supporting data that feeds AI agents and retrieval systems, not only dashboards and reports.
Practical notes
This is a fully remote role supporting a team in the EST time zone (9 AM to 5 PM EST).
The position is full-time and based in Pakistan with no relocation required.
You will work closely with an experienced team that operates at real marketing scale, using real customer data from day one.