Staff Data Engineer
hims-and-hersRemote (USA)Full Time6d ago
PythonGoAWSGCPTerraformCI/CDMLData ScienceETLAirflowKafkaSecurity
Job description
Staff Data Engineer at hims-and-hers
About the role
Join the Data Platform Engineering team at Hims & Hers as a key technical leader. You will influence architectural choices, improve system reliability, and enhance the developer experience for a team of nine engineers building the core infrastructure that supports millions of users.
Key facts
What you'll do
- Lead complex, multi-sprint platform initiatives including Fivetran connector development, Databricks Lakehouse migration, event streaming infrastructure, and engineering standards adoption.
- Design, construct, and maintain production-ready data ingestion pipelines and platform infrastructure, serving as the foundation for Analytics Engineering, Data Science, and business teams.
- Build and operate event-driven and streaming data pipelines using Kafka, PySpark, and Databricks Structured Streaming, defining scaling, cost controls, and monitoring strategies.
- Manage data contracts, schema governance, and service level agreements for the ingestion and raw-to-cleansed data layers.
- Ensure data quality for developed pipelines through dbt tests, anomaly detection, schema validation, and data drift alerting.
- Maintain system reliability by defining key performance indicators and service level objectives, implementing monitoring as code, and participating in the on-call rotation.
- Oversee the integration and data activation layers, including Fivetran connectors and Hightouch reverse ETL connectors, from infrastructure provisioning to production monitoring.
- Support Analytics Engineers, Data Scientists, and ML engineers by building platform capabilities and pipelines that unblock their work.
- Identify and address systemic inefficiencies within the Data Platform Engineering team's pipelines and infrastructure.
- Mentor Senior Data Engineers through design and code reviews, fostering their growth into cross-squad responsibilities.
- Drive the adoption of engineering best practices, including testing, CI/CD, observability, and schema registry governance.
Requirements
- A minimum of 8 years of professional experience in designing, building, and operating data pipelines and platform infrastructure.
- Demonstrated experience with Change Data Capture (CDC) patterns for real-time data ingestion.
- Experience with Flink for stream processing.
- Proven ability to govern and administer dbt in production BigQuery or Databricks environments, including CI/CD, testing standards, documentation, and schema governance for ingestion-layer pipelines.
- Experience building and managing Airflow DAGs at scale, focusing on task orchestration, reliability, and scheduling.
- Experience building event streaming pipelines with Kafka or Confluent Kafka, including producer/consumer management, schema evolution, Schema Registry, and consumer lag monitoring.
- Proficiency in both GCP and AWS cloud environments, with daily use of BigQuery on GCP and Airflow on AWS EKS.
- Experience owning data quality for production pipelines, including dbt tests, anomaly detection, and alerting on schema changes and data drift.
- Experience with Fivetran or similar connector platforms, covering infrastructure as code provisioning, schema change management, and health monitoring.
- Experience with the Databricks platform, including Delta Lake, Databricks Workflows, and Unity Catalog.
- Familiarity with data compliance in regulated environments, such as HIPAA/PHI handling, access controls, and audit logging.
- Experience with infrastructure as code using Terraform or comparable tools.
- Strong proficiency in Python and SQL for writing and reviewing production-grade pipeline code.
- Excellent design skills, with the ability to translate ambiguous requirements into clear solutions and deliver them to production with minimal revisions.
Nice to have
- Experience with PySpark/SparkSQL for large-scale data processing.
- Experience with Hightouch or similar reverse ETL platforms.
- Familiarity with MLOps, supporting ML engineers with data pipelines for model training or feature stores.
- Familiarity with Looker LookML or equivalent BI serving layers.
- Go experience for Kafka service development.
- Experience in direct-to-consumer healthcare, telehealth, or similarly regulated industries.
- Familiarity with UK/GDPR data compliance requirements.
Skills & tools
- BigQuery
- dbt
- Airflow on Astronomer
- Confluent Kafka
- Databricks
- Fivetran
- Terraform/OpenTofu
- Python
- SQL
- GCP
- AWS
- Change Data Capture (CDC)
- Flink
- PySpark
- SparkSQL
- Hightouch
- MLOps
- Looker LookML
- Go
Practical notes
- This is a full-time, remote position based in the US.
- Hims & Hers offers competitive salary and equity, unlimited PTO, company holidays, quarterly mental health days, comprehensive health benefits, ESPP, and 401k with employer matching.
- The company is committed to building a diverse workforce and encourages applications from candidates who may not perfectly match every requirement.
- Hims & Hers complies with fair chance ordinances and laws in California, San Francisco, and Los Angeles.
- In Massachusetts, lie detector tests are unlawful as a condition of employment.
- Reasonable accommodations are available for qualified individuals with disabilities during the application process. Please contact accommodations@forhims.com for assistance.