Staff Data Engineer
Job description
About the role
You will architect the platform by setting our warehouse and lakehouse direction while establishing secure, query-ready foundations for payments data. You will design and operate batch and streaming workflows that reliably move information from source systems into governed storage using change data capture. You will establish canonical models and clear contracts so the entire organization understands dimensions, grains, and semantics of financial information. You will implement rigorous quality checks, observability dashboards, and monitoring for financial datasets to ensure correctness is built in from the start. In a regulated environment, you will implement role-based protections, masking, and lineage tracking to keep sensitive information secure from day one. You will construct feature pipelines and reliable tables that enable analytics teams and machine learning workflows to operate at scale. This role requires partnership with platform and product teams to ensure data infrastructure supports critical monetization and integration efforts effectively. The choices you make in your first quarter will still be load-bearing years from future, defining the path for every data engineer who follows you.
Key facts
What you'll do
- Architect the platform by defining warehouse and lakehouse direction and establishing secure foundations for payments data.
- Design and operate batch and streaming workflows that reliably move information from source systems into governed storage using change data capture.
- Establish canonical models and clear contracts so the entire organization understands dimensions, grains, and semantics of financial information.
- Implement rigorous quality checks, observability dashboards, and monitoring for financial datasets to ensure correctness and reliability.
- Implement role-based protections, masking, and lineage tracking to keep sensitive financial information secure from day one in a regulated environment.
- Construct feature pipelines and reliable tables that enable analytics teams and machine learning workflows to operate effectively at scale.
- Define practices, tooling, and deployment standards that future engineers can inherit and improve upon over time.
- Partner with platform and product teams to ensure data infrastructure supports critical monetization and integration efforts.
- Mentor colleagues and raise the technical bar by example through code and design reviews.
- Champion data correctness through observability and rigorous quality checks across all financial datasets.
- Operate within major cloud environments such as AWS, GCP, or Azure, applying knowledge of security and cost management.
- Handle sensitive or regulated data with appropriate access controls, encryption, and governance to limit organizational risk.
Requirements
- Bring eight or more years of experience building production data systems with a track record of owning architecture and seeing big decisions through to production.
- Write expert SQL and proficient Python to solve complex data integration and transformation challenges in production environments.
- Possess deep experience in at least one modern lakehouse or warehouse ecosystem, such as Snowflake or Databricks, and reason across different technologies rather than relying on specific product knowledge.
- Understand dimensional modeling, normalized structures, or Data Vault approaches and design models that remain useful as systems evolve.
- Orchestrate large-scale pipelines using tools such as Airflow, Dagster, or Prefect in production settings with reliability and scalability.
- Operate comfortably on major clouds including AWS, GCP, or Azure, applying knowledge of security and cost management best practices.
- Have handled sensitive or regulated data, implementing access controls, encryption, and governance to protect financial and personal information.
- Raise the technical bar by example, mentoring colleagues and improving both code and design reviews consistently.
Nice to have
- Background in payments, fintech, or regulated domains, including familiarity with PCI DSS controls and tokenization approaches.
- Experience with streaming platforms such as Kafka or managed alternatives, and streaming processing frameworks for real-time data flows.
- Hands-on experience with data governance, lineage visualization, and observability tools including Unity Catalog or similar systems.
- Experience supporting machine learning, feature stores, and pipelines used for training and inference in production.
- An interest in mentoring peers, contributing to open source data practices, and growing platform capabilities over time within the organization.
Practical notes
The role is remote with no required travel; visa sponsorship is not available at this time. Standard working hours apply.