Lead Data Engineer Lille
Job description
About the role
The is responsible for directing and elevating the data engineering capabilities of client projects across diverse sectors and company sizes. This role owns the full lifecycle of data pipeline development, from initial ingestion and complex processing to secure and efficient data exposure for business consumption. The hire will act as the primary technical liaison, ensuring seamless communication and alignment between IT departments and business stakeholders such as business experts, data scientists, and data analysts. They will enforce rigorous standards for data architecture by validating design choices and overseeing the transition from development to stable production environments. Furthermore, this position demands active participation in the industrialization of data science models, focusing on scalability, testing, and maintaining high levels of craftsmanship. The role is deeply embedded within a collaborative community that values knowledge exchange, where technical challenges are openly discussed and refined through internal and external events. Success in this position is measured by the ability to deliver robust, scalable data solutions while simultaneously mentoring the team and contributing to the collective expertise of the Ippon data community.
Key facts
What you'll do
Orchestrate the development and deployment of end-to-end data ingestion, transformation, and exposure pipelines using modern distributed frameworks.
Evaluate and validate complex data architecture designs to ensure they meet functional, non-functional, and scalability requirements before implementation.
Oversee the operationalization, monitoring, and long-term maintenance of production data flows to guarantee reliability and performance.
Collaborate closely with data science teams to provide the infrastructure and support necessary for industrializing machine learning models and analytics workloads.
Implement infrastructure as code solutions using tools like Terraform and CloudFormation to provision and manage cloud-based data processing environments.
Champion active participation in internal and external data community events to foster knowledge sharing and professional growth.
Contribute thought leadership to the data community by authoring blog posts, documenting return-on-investment cases, and preparing internal and external conference talks.
Leverage practical experience from numerous client missions to mentor junior engineers and elevate the overall craftsmanship of the data platform.
Ensure that all delivered solutions adhere to the operational demands of production delivery within fast-paced, agile client environments.
Drive the adoption of streaming data technologies to support real-time processing and immediate business insights.
Establish best practices for data quality, governance, and security embedded within the pipeline development lifecycle.
Translate evolving business requirements into robust technical specifications that guide the data engineering team effectively.
Continuously explore emerging data technologies and cloud services to identify opportunities for improving the current architecture and tooling.
Act as a senior technical resource during discovery phases to define feasible and high-impact data strategies for clients.
Requirements
You hold a Bac+5 degree or equivalent higher education with a first professional experience focused on data engineering or a related field.
You demonstrate expert-level proficiency with at least one major distributed computing framework such as Spark, Storm, or Flink.
You possess strong programming skills in at least one of the following languages: Python, Java, C/C++, or Scala.
You have hands-on experience with a wide variety of database systems, including both SQL and NoSQL databases, and are highly competent in writing complex SQL queries.
You have practical experience with data streaming platforms, specifically Kafka, Kinesis, or similar technologies.
You have a proven track record working with major cloud platforms such as AWS, GCP, or Azure, managing infrastructure and data services.
You have successfully designed and deployed multiple data architectures from scratch, guiding them through the entire lifecycle until they reach production stability.
You are accustomed to working within agile methodologies and are capable of delivering high-quality code and solutions in demanding production settings.
You have a deep understanding of the challenges involved in deploying data solutions at scale, including considerations for performance, cost, and maintainability.
You are passionate about sharing knowledge and actively contributing to the technical growth of peers through internal discussions and community events.
You are comfortable working remotely and managing your schedule to meet deadlines without direct supervision.
You have a strong sense of ownership regarding the quality and reliability of the data pipelines you oversee.
You are able to communicate effectively with both technical and non-technical stakeholders to align on objectives and constraints.
You have a history of continuous learning, constantly updating your skills to keep pace with the rapidly evolving data landscape.
Nice to have
Experience with data visualization tools to better communicate insights to business teams.
Knowledge of data governance frameworks and compliance standards relevant to regulated industries.
Familiarity with containerization and orchestration technologies such as Docker and Kubernetes.
Practical notes
This is a permanent contract position based on remote work.
Candidates must be available to work within the standard French business hours to ensure effective collaboration with European clients and teams.