Senior Data Engineer
Job description
Senior Data Engineer at Veeva.
About the role
Join the NitroAI team to build and maintain the core data infrastructure that drives analytics. This role involves close collaboration with data scientists and other analytics professionals to transition manual processes into production, establish governed data pipelines, and deliver data efficiently to customers. It's a high-impact position within a startup-like environment at Veeva, offering significant influence over data delivery across a growing multi-tenant platform.
Key facts
What you'll do
- Manage the Airflow codebase, including approximately 100 DAGs on MWAA, by creating reusable templates, scaling patterns, enforcing standards, and improving testing.
- Act as the primary resource for delivery teams regarding pipeline architecture and troubleshooting.
- Oversee data onboarding for new sources, which includes setting up connections, discovering schemas, and designing initial pipelines.
- Operate and expand large-scale Spark pipelines on AWS Glue, handling multi-terabyte joins and compaction, and support their migration to Databricks.
- Enhance data models for commercial pharmaceutical data, focusing on patient claims, KOL/HCP data, and CRM activity, with an emphasis on structure, lineage, and data flow.
- Contribute to and maintain the internal Python package utilized by the data team.
Requirements
- 5 or more years of experience in building data models and pipelines.
- 2 or more years of experience with production Airflow.
- 2 or more years of experience with Spark at scale, specifically with multi-terabyte datasets; AWS Glue experience is preferred.
- Strong proficiency in Python and SQL.
- Fluent in AWS services such as S3, ECS/Fargate, Glue, IAM, and Secrets Manager.
- Excellent communication skills and a capability to teach, as onboarding analytics teams to the codebase and developing their Airflow skills is a key responsibility.
Nice to have
- Experience with privacy-sensitive or governed data.
- Databricks experience, particularly with migrating Glue workloads.
- Proficiency with Claude Code.
- Familiarity with data science workflows and ML pipeline tools.
- A background in life sciences or healthcare.
- Experience with Open Meta data or similar data catalog and dictionary tools.
Skills & tools
- Airflow (MWAA)
- Spark (AWS Glue, Databricks)
- Python
- SQL
- AWS (S3, ECS/Fargate, Glue, IAM, Secrets Manager)
- Claude Code (nice to have)
- Open Meta data (nice to have)
Practical notes
This is a Work Anywhere company, supporting flexibility for remote or office work. The base salary range for this role is $115,000 to $175,000, with potential for variable bonus and/or stock bonus. Benefits include medical, dental, vision, basic life insurance, flexible PTO, company paid holidays, retirement programs, and a 1% charitable giving program. The application process involves submitting a resume, completing a personality assessment within 3 days, and then potentially proceeding to interviews with the hiring manager, a practical case exercise, and a final conversation with a Senior Leader.