Stage - Data Governance Engineer H/F/X
Job description
About the role
As the Data Governance Engineer intern at Veepee, you will own the design and implementation of intake pipelines that unify fragmented data sources across the business. You will enforce data governance contracts to ensure compliance, accuracy, and reliability of information assets used by analytics and science teams. This role requires you to build robust foundations for data consumption while maintaining strict quality standards across all datasets. You will collaborate closely with technical and analytical partners to translate governance requirements into technical solutions. You will implement transformation logic using dbt to support flexible and reliable reporting environments for multiple departments. Your work will ensure that data is structured, documented, and accessible for downstream users and external sharing initiatives. You will oversee data catalog clarity and manage definitions to maintain consistency across publishing layers. Finally, you will coordinate validation efforts with technology partners to guarantee stable, performant, and traceable production deployments.
Key facts
What you'll do
Orchestrate migration workflows using Apache Airflow to transition towards on-premise setups while managing dependencies and scheduling complexities.
Implement transformation logic through dbt models to support analysts and data scientists across various departments with reliable and flexible data structures.
Design and develop intake pipelines that consolidate diverse data sources to feed analytics, science, and external sharing initiatives efficiently.
Champion data quality by defining, monitoring, and enforcing rules and governance contracts that protect the integrity and compliance of information assets.
Construct dbt artifacts aligned with metadata policies to support analytics consumption and reporting needs while adhering to established standards.
Utilize Trino to execute queries and interact with the Trino and Iceberg Metastore environment for performant data access and management.
Oversee data catalogs to ensure clarity and rule enforcement for datasets shared through publishing layers, including managing definitions and access controls.
Coordinate with technology partners to validate pipelines before production deployment, verifying documentation completeness and ensuring traceable lineage and usage expectations.
Monitor and optimize data flows within Google Cloud Dataflow and Google Cloud Pub/Sub to maintain reliable and scalable ingestion and messaging processes.
Leverage Google BigQuery to store and process structured data, ensuring efficient querying and integration within the broader data ecosystem.
Apply data governance frameworks to guide daily work, ensuring that all data activities respect policies, standards, and business requirements.
Collaborate with analysts and data scientists to understand their data needs and translate them into governed, self-serve data assets and reporting layers.
Document data flows, transformations, and governance rules to provide clear references for stakeholders and ensure long-term maintainability.
Act as a bridge between technical implementation and business expectations to ensure that data initiatives deliver tangible value and trust.
Requirements
You are currently studying computer science or engineering with a data specialization, bringing a strong academic background and intellectual curiosity to the role.
You write development code and tests with rigor, applying best practices to ensure maintainable, scalable, and reliable solutions.
You think in agile patterns, working effectively in team settings and embracing collaboration with peers, product owners, and technical leads.
You possess a deep understanding of SQL, using it to query, transform, and shape information assets for diverse analytical needs.
You know Python for scripting and process automation tasks, leveraging libraries and structures to improve efficiency and reduce manual effort.
You respect data contracts and governance definitions in daily work, adhering to standards and ensuring consistency across all data operations.
You demonstrate ownership and accountability for data quality, proactively identifying issues and driving improvements across pipelines and datasets.
You communicate clearly and precisely, both in writing and verbally, to articulate technical concepts to non-technical stakeholders and align on objectives.
Nice to have
No additional preferred items are specified for this role.
Practical notes
This internship is full-time and located in Saint-Denis.
The monthly compensation for this internship is fixed at 16,000 euros.
The engagement is time-bound, and candidates must align their availability with the internship schedule.
Travel requirements are not specified, but the role operates within the defined Saint-Denis location.
There are no visa requirements mentioned for this position, as it is governed by standard internship regulations in France.
Candidates must ensure they meet the academic and technical criteria outlined in the requirements section to be considered for this opportunity.
Applications must respect the framework defined by the hiring team, focusing on technical alignment and demonstrated capability in data governance and engineering practices.