Data Engineer
Job description
About the role
At Preply, we are reimagining education through a human-led, tech-enabled approach that creates life-changing learning experiences. Our recent $150M Series D has accelerated our vision to transform education at global scale, reaching 100,000+ tutors teaching 90+ languages to learners in 180 countries. As a category-defining unicorn, we are building the future of learning and need dedicated talent to support this mission. The Data Ingestion and Enrichment team plays a critical role by providing a single, trusted, and scalable data foundation. This team ensures that analytics, machine learning, and product features are built on unified, governed, and production-grade data assets within Preply's Lake House. This includes extracting, normalizing, and generating structured data from unstructured assets to form a durable data moat for AI-driven products.
Joining Preply as a Data Engineer means helping define the future of education at global scale while building something that truly matters for millions of people every day.
What you'll do
You will own the design, implementation, and operation of the core data ingestion and enrichment platforms that power Preply's learning marketplace. This role is central to ensuring every analytics dashboard, machine learning model, and product feature rests on a foundation of trusted, well-documented, and observable data. You will define and enforce data contracts, schemas, and quality standards so downstream teams can consume data with confidence. Collaboration with ML Platform, Data Science, Analytics Engineering, and Product teams will turn business questions into robust, production-grade data capabilities. You will instrument pipelines for latency, freshness, and cost to drive reliability and proactive operations. Additionally, you will enable self-service by creating reusable templates, shared libraries, and clear documentation so new data sources can be onboarded quickly and safely.
- Designing, building, and maintaining the data lake infrastructure and data as a product foundations that support analytics, machine learning, and product decisions at scale.
- Delivering end-to-end batch and streaming ingestion pipelines that balance performance, cost, and reliability while defining clear raw, standardized, and consumption layers with explicit lineage and retention policies.
- Implementing data contracts and quality checks, including schema validation, freshness, volume, and anomaly detection, to prevent bad data from reaching consumers.
- Building enrichment and cross-domain join logic that standardizes and contextualizes data, ensuring historical tracking, point-in-time correctness, and safe dataset versioning.
- Instrumenting pipelines with comprehensive observability for freshness, latency, data quality, and cost, and contributing to SLOs, alerting, and incident response processes.
- Applying access control, classification, and privacy protections at ingestion time to ensure sensitive data is masked, minimized, or anonymized and that all data flows remain auditable.
- Creating standardized ingestion templates, shared libraries, and platform tooling that empower teams to onboard new data sources independently with high discoverability and metadata clarity.
- Collaborating with Product, Backend, Analytics, and ML partners to align on requirements and trade-offs, while mentoring peers and junior engineers to uphold data quality standards and data contracts.
Requirements
To succeed in this role, you will need hands-on experience building components of large, high-scale applications, such as data pipelines, well-structured APIs, and efficient algorithms that handle substantial data volumes. You should have a solid background working within platform or data engineering teams in multi-stakeholder environments where delivery depends on cross-functional alignment. Demonstrated familiarity with cloud platforms like AWS or GCP and modern DevOps practices, including infrastructure as code, CI/CD, and containerization, is essential.
You must possess hands-on experience designing and implementing real-time and batch data processing systems, including stream processing, windowing, and stateful operations. Strong proficiency in at least one modern data stack, including data lake frameworks, workflow orchestration tools, and data quality platforms, is required. Experience defining and enforcing data contracts, with a track record of implementing schema evolution, freshness checks, volume monitoring, and data quality rules, is necessary. Deep understanding of data modeling for analytics, including dimensional modeling, slowly changing dimensions, and techniques for ensuring point-in-time correctness, is expected. Commitment to operational excellence, including instrumenting pipelines, setting up monitoring, and participating in on-call and incident response practices, is mandatory.
Nice to have qualifications include preferred experience working with data quality and observability tools that support anomaly detection and lineage, as well as familiarity with data catalog and metadata discovery tools that improve dataset discoverability and trust.
Practical notes: The engagement is full_time, and the position is located in Barcelona.