
Senior Data Engineer
Job description
About the role
You will serve as the definitive subject matter expert for our next-generation Databricks Lakehouse platform, guiding its architecture and long-term technical vision. In this capacity, you will own the design and construction of scalable ingestion pipelines that handle diverse and evolving data demands. You will establish and enforce production-grade best practices that raise the engineering standard across the entire data team. Your work will set the CI/CD standards for data pipelines, ensuring that deployments are reliable, automated, and auditable. You will tackle complex problems involving unstructured and semi-structured data sources with a pragmatic, solution-oriented mindset. You will mentor and elevate the capabilities of your peers through clear communication and hands-on technical leadership. This role requires a self-directed engineer who thrives in a fast-paced rebuild environment and can drive decisive architectural standards.
Key facts
What you'll do
Design, implement, and maintain robust Lakehouse ingestion and transformation pipelines using Databricks and PySpark to meet evolving business needs.
Establish and enforce platform standards and best practices that level up the engineering maturity of the broader data team.
Set up automated delivery pipelines utilizing Databricks Asset Bundles (DABs) and modern orchestration tools such as Airflow to ensure seamless and reliable deployments.
Architect and deliver pipelines that ingest, process, and structure both structured online transaction processing data and unstructured or semi-structured data from APIs, logs, and cloud storage streams.
Enforce Delta Lake principles including change data capture, schema evolution, and optimization to maintain a high-quality data foundation.
Integrate data quality frameworks like Great Expectations and robust governance through Unity Catalog into day-to-day development and deployment workflows.
Direct, evaluate, and refine the use of AI agents for rapid code generation, ensuring all generated code meets strict standards for quality, security, and performance.
Communicate complex technical trade-offs clearly and directly to both technical and non-technical stakeholders, aligning solutions with business objectives.
Mentor team members by articulating best practices and driving knowledge transfer to improve overall engineering capability.
Drive pragmatic, autonomous solutions for ambiguous and legacy datasets without needing step-by-step guidance.
Confidently lead technical decisions and implementation details in a fast-paced environment where requirements evolve quickly.
Leverage deep expertise in Spark and Databricks to optimize distributed workloads and resolve performance bottlenecks at scale.
Partner with cross-functional teams to translate business requirements into durable data architectures and pipeline specifications.
Continuously assess and introduce new tools, patterns, and frameworks that enhance the reliability and efficiency of the Lakehouse platform.
Requirements
Possess 5 or more years of dedicated data engineering experience with a strong focus on production Databricks and PySpark at scale.
Have extensive, proven experience building large-scale data pipelines in production environments using Databricks and Apache Spark.
Demonstrate strong proficiency across major cloud ecosystems, with AWS as the primary environment and equal comfort with Azure or GCP where Databricks is deployed.
Have hands-on, enterprise-level experience implementing CI/CD pipelines for Databricks, specifically using Databricks Asset Bundles or comparable infrastructure-as-code methodologies.
Show a documented history of processing raw, unstructured, or semi-structured data sources into structured, production-ready layers within a Lakehouse architecture.
Exhibit expert-level SQL optimization skills combined with deep Python proficiency for distributed data processing.
Have practical, real-world experience with workflow orchestration tools such as Airflow and data governance frameworks including Unity Catalog and data quality tools like Great Expectations.
Be a senior-level communicator capable of clearly explaining intricate technical concepts to diverse audiences and mentoring peers on engineering best practices.
Embrace a pragmatic, autonomous approach to solving complex data engineering challenges in ambiguous and legacy-heavy contexts.
Comfortably guide the evaluation and implementation of AI-assisted coding tools while maintaining rigorous standards for code quality and stability.
Nice to have
Experience contributing to or building internal platform engineering tools that support data teams.
Background in industries with strict compliance or regulatory requirements.
Practical notes
This is a full-time remote position with flexible hybrid working arrangements.
The role may involve occasional travel as needed for team collaboration or client engagements.
Candidates must be eligible to work in the country where the position is based.
No specific deadline for applications is provided in the source material.