Data Engineer
Job description
About the role
You are the data architect and steward for the Safety and Customer Care (SCC) ecosystem, owning the design and execution of the data models and pipelines that power over 1.7 million monthly human and AI interactions. You ensure the reliability and performance of critical data infrastructure that directly impacts rider and driver safety experiences. Your work transforms support interactions into actionable insights, fueling AI agents and human associates to deliver genuine connection. You will implement systems that track data quality and consistency while bridging business goals with technical execution. Your role is pivotal in saving millions of dollars annually through efficiency improvements and accurate reporting. You will collaborate cross-functionally to unlock insights from massive Lyft data sets that fuel Analytics, Data Science, Engineering, and Marketing teams.
Key facts
What you'll do
Develop and own the core data pipeline architecture to scale data processing flow in response to rapid data growth at Lyft.
Evolve data models and schemas in partnership with business and engineering stakeholders to align with changing requirements.
Implement robust systems for tracking data quality and consistency across SCC data assets.
Build and maintain self-service data pipeline management tools to enable ETL efficiency and team autonomy.
Tune SQL and MapReduce jobs to optimize data processing performance and reduce computational costs.
Write clean, well-tested, readable, and maintainable code that sets a high standard for engineering excellence.
Participate in code reviews to ensure quality, distribute knowledge, and drive best practices across the team.
Collaborate cross-functionally with product, engineering, data science, and marketing teams to deeply understand business problems and co-own solutions.
Leverage huge volumes of Lyft data to generate insights that improve safety, support efficiency, and user satisfaction.
Serve as a technical partner for AI agent platforms, ensuring data reliability and feature correctness for critical workflows.
Optimize data storage and retrieval strategies to support real-time and near-real-time decision-making needs.
Act as a subject matter expert on data infrastructure, guiding associates on best practices and scalable design patterns.
Champion data governance initiatives to ensure compliance, security, and consistency across all SCC data products.
Drive continuous improvement by monitoring pipeline health, identifying bottlenecks, and implementing resilient solutions.
Requirements
Hold a Bachelor's degree in Computer Science, Engineering, Mathematics, Statistics, or a related field.
Bring 4+ years of professional experience in data engineering, ideally with large-scale distributed systems in production environments.
Demonstrate strong skills in Spark, Python (or similar scripting language), and SQL performance tuning for complex queries.
Show experience with AWS, Hadoop/S3, Hive, Presto, Airflow, and related data infrastructure tools.
Possess a solid understanding of ETL processes, workflow orchestration, and data warehousing concepts and practices.
Exhibit a collaborative mindset, thriving when working across teams to solve real-world and ambiguous problems.
Commit to writing unit tests and integration tests to ensure data pipeline reliability and correctness.
Adhere to version control best practices and documentation standards to support maintainable and reproducible pipelines.
Maintain a strong sense of ownership for data quality, accuracy, and timeliness across all SCC data products.
Embrace security and privacy principles, ensuring data handling aligns with company policies and regulatory requirements.
Communicate technical concepts clearly to both technical and non-technical stakeholders, including product managers and business partners.
Willingness to rotate on-call responsibilities for critical data pipelines and respond promptly to incidents.
Adaptability to changing priorities and ability to manage multiple competing deadlines in a fast-paced environment.
Commitment to continuous learning and staying current with data engineering trends, tools, and technologies.
Nice to have
Experience with building and maintaining data observability and monitoring frameworks.
Familiarity with machine learning data pipelines and feature stores used for AI model training and inference.
Knowledge of containerization and orchestration tools such as Docker and Kubernetes for data workloads.
Understanding of privacy-preserving data practices and anonymization techniques for sensitive user data.
Experience contributing to open source data engineering projects or publishing technical content.
Practical notes
This role requires in-office presence in Toronto, Canada, to foster collaboration and innovation.
Travel is not expected as part of the standard work arrangement for this position.
Visa sponsorship may be considered for eligible candidates based on role requirements and policy.
Employment is contingent upon verification of identity and authorization to work in Canada.
Candidates must comply with Lyft's background check and security policies as a condition of employment.