Senior Data Engineer
Job description
About the role
Proton is constructing the foundational data layer for an AI operating system that fundamentally reshapes wholesale distribution. This role is tailored for a builder who wants to own and evolve the core data infrastructure that powers intelligent systems. You will work closely with AI agents to deliver high-quality, trustworthy data that serves as the backbone for critical business decisions. The position demands a high degree of ownership over the entire data lifecycle, from initial ingestion to final consumption by end-users. You will be responsible for ensuring that data is reliable, performant, and secure at every stage of its journey. Collaboration with AI coding assistants is an integral part of the daily workflow, used to build and ship production-ready code efficiently. The focus will be on establishing robust systems for data validation, reconciliation, lineage tracking, and error handling to guarantee data accuracy. Finally, you will partner closely with engineering and product teams to define and enforce clear data agreements that align technology with business objectives.
Key facts
What you'll do
- Assume full responsibility for the data layer, covering the complete spectrum of ingesting diverse data sources, transforming them, and serving data to both products and AI applications.
- Develop, maintain, and optimize data pipelines and transformations using modern cloud data warehousing and orchestration tools to ensure scalability and reliability.
- Integrate and reconcile large, complex datasets originating from a wide variety of sources, including files, event streams, and external APIs, into a coherent structure.
- Design and implement multi-layered data models across distinct layers such as raw, refined, and curated to guarantee data reliability, usability, and long-term maintainability.
- Contribute proactively to the future architecture and expanded capabilities of the data layer, influencing strategic technical decisions that shape the platform.
- Utilize AI coding assistants on a daily basis to write, debug, and refine production-ready code, placing strong emphasis on prompt engineering techniques and rigorous output validation.
- Establish comprehensive systems for data validation, reconciliation, data lineage, and error handling to guarantee the highest levels of data accuracy and trustworthiness.
- Collaborate closely with engineering and product teams to define, document, and maintain clear data agreements that align expectations and ensure consistency.
- Optimize data performance and cost-efficiency within the cloud data warehouse, balancing query speed with resource utilization.
- Implement monitoring and alerting frameworks to detect data quality issues, pipeline failures, or anomalies in real-time, enabling swift remediation.
- Lead documentation efforts for data schemas, pipelines, and processes to ensure knowledge transfer and operational continuity.
- Partner with data scientists and analysts to ensure that the data layer supports advanced analytics, machine learning, and business intelligence initiatives.
- Evaluate new tools, frameworks, and methodologies to continuously improve the data stack and incorporate industry best practices.
- Act as a technical leader in data engineering, mentoring junior team members and elevating the overall standard of code and design.
Requirements
- Possess a minimum of 7 years of professional experience as a data engineer, with demonstrable ownership of production data systems and models in a live environment.
- Exhibit a strong and deep understanding of core data engineering principles, including query optimization strategies and pipeline debugging techniques across distributed systems.
- Demonstrate high proficiency in programming and SQL to build efficient, scalable data pipelines and robust database schemas.
- Have substantial experience building reliable ingestion and transformation pipelines that handle varied data sources, formats, and schemas.
- Hold hands-on experience with at least one major cloud provider and their corresponding cloud data warehouse solutions, such as Snowflake or similar platforms.
- Show proven ability to ingest data from file-based sources, event or streaming sources, and RESTful APIs, managing the complexities of each.
- Maintain a solid knowledge of common data consistency issues, such as race conditions and partial updates, and effectively apply mitigation strategies.
- Use AI development tools like Claude Code or similar platforms for software development on a daily basis, integrating them into your standard workflow.
- Illustrate a consistent track record of taking data systems from the conceptual phase through to production deployment, applying strong technical judgment and decision-making.
- Exhibit a pragmatic, fast-paced working style that thrives in a startup environment, balancing speed with quality and sustainability.
- Communicate effectively in writing, ensuring clarity, precision, and professionalism in all documentation and correspondence.
- Achieve a minimum English language proficiency level of C1 or higher, enabling fluent participation in all professional interactions.
Nice to have
- Extensive, hands-on experience with cloud data warehouses and modern data transformation tools, pushing the boundaries of what is possible.
- Deep experience with large-scale streaming or event-based data ingestion using technologies such as Kafka or Pulsar.
- Familiarity with medallion or lakehouse architectures for managing complex data environments with strict governance requirements.
- Prior experience integrating with large enterprise systems such as ERPs or e-commerce platforms, navigating their intricacies.
- Experience building data systems specifically designed for AI and machine learning applications, understanding the unique demands of these workloads.
- Previous work experience within an early-stage SaaS startup, where agility and ownership are highly valued.
Practical notes
- Requires meaningful daytime overlap with the Boston (EST) team.
- Visa sponsorship is not provided.