Principal Data Engineer
hims-and-hersRemote (USA)Full Time5d ago
PythonGoAWSGCPTerraformCI/CDMLData EngineeringData ScienceETLAirflowSpark
Job description
Principal Data Engineer at hims-and-hers
About the role
This is a senior individual contributor position focused on shaping the future of our data platform. You will be instrumental in defining and executing the architectural direction for data ingestion, processing, and serving, aligning closely with product and engineering leadership to meet business objectives. Your influence will span the entire organization, setting technical standards and guiding critical architectural choices.
Key facts
What you'll do
- Guide the long-term technical blueprint for data platform engineering, covering data intake, workflow management, real-time event streams, and the underlying infrastructure for data transformation and delivery.
- Lead architectural reviews and serve as the primary technical point person for significant changes impacting multiple systems and costs.
- Establish and enforce engineering best practices and production readiness standards for all data platform systems, including testing, deployment pipelines, observability, logging, data contracts, and schema management.
- Design and implement data quality and observability frameworks, such as anomaly detection, schema validation, and data drift alerting, to ensure data trustworthiness.
- Define the strategy for self-service analytics capabilities, empowering analytics engineers while establishing guardrails to prevent downstream issues.
- Oversee the selection, integration, and ongoing management of data platform tools, including contract negotiation, cost monitoring, and retirement decisions.
- Develop and manage data sharing mechanisms, access controls, cross-team data agreements, and governed data consumption pathways.
- Drive cost optimization for platform infrastructure, including BigQuery resource management, query efficiency, and cloud spend accountability.
- Lead incident response for critical platform issues, acting as a technical escalation point and facilitating post-mortems to prevent recurrence.
- Produce high-quality technical documentation, such as architecture decision records and design proposals, to foster alignment and serve as team references.
- Mentor and develop Staff and Senior Data Engineers, elevating the team's technical capabilities through reviews and hands-on collaboration.
- Collaborate with Machine Learning, Data Science, Legal, Security, and DevOps teams to deliver compliant and robust platform features.
- Contribute directly to critical development tasks.
Requirements
- Over 15 years of professional experience in designing, building, and managing large-scale data platform architectures.
- Proven ability to align architectural vision with organizational business objectives.
- Extensive expertise in cloud-native data platforms, with a strong preference for GCP (BigQuery, GCS, Dataflow) and operational familiarity with AWS (EKS-based Airflow).
- Hands-on experience with modern data stack components, including dbt at scale, Airflow/Astronomer, Kafka/Confluent, Databricks/Spark, Fivetran, and data activation platforms like Hightouch.
- Experience designing and operating event streaming pipelines, including Schema Registry, data contracts, and consumer lag management.
- Demonstrated success in establishing engineering standards across multiple teams and driving adoption without direct authority.
- Experience owning data quality frameworks, including dbt testing, anomaly detection, and data observability tools.
- Experience with data governance and compliance in regulated environments, such as HIPAA/PHI handling, data classification, and access controls.
- Experience leading incident response for data platform outages, including root cause analysis and operational improvements.
- Proficiency in infrastructure-as-code tools like Terraform.
- Strong proficiency in Python and SQL, with the ability to write and review production-grade code.
- Excellent written communication skills and the ability to operate effectively in ambiguous situations.
Nice to have
- Experience with Databricks, Unity Catalog, and Delta Lake in production environments.
- Experience with Change Data Capture (CDC) patterns and Apache Flink for real-time processing.
- Experience with PySpark/SparkSQL for large-scale batch and streaming workloads.
- Experience managing migrations from BigQuery to Databricks Lakehouse or similar cloud data warehouse migrations.
- Experience with MLOps, including model training pipelines and feature stores.
- Experience developing Kafka producers and consumers using Go or Python.
- Experience in a direct-to-consumer healthcare or telehealth company with HIPAA and GDPR compliance.
- Familiarity with SOX compliance controls in a data engineering context.
Skills & tools
- GCP (BigQuery, GCS, Dataflow)
- AWS (EKS)
- Airflow/Astronomer
- dbt
- Kafka/Confluent
- Databricks/Spark
- Fivetran
- Hightouch
- Terraform
- Python
- SQL
- Apache Flink
- PySpark/SparkSQL
- Go
Practical notes
- This role offers competitive salary and equity.
- Benefits include unlimited PTO, company holidays, mental health days, comprehensive health coverage, ESPP, and 401k with employer match.
- The company is committed to building a diverse workforce and encourages applications from candidates who may not perfectly match every requirement.
- Hims & Hers complies with fair chance employment ordinances and laws.
- Accommodation is available for individuals with disabilities during the application process.