Senior Data Engineer
Job description
About the role
Join the Data Foundations team to build and scale the core systems behind the Healthcare Map. You will transform complex healthcare data into reliable assets that drive internal innovation and customer-facing applications. In this role, you own the design, implementation, and reliability of large-scale data infrastructure that powers critical insights across the organization. You will be responsible for developing robust pipelines that ingest, process, and serve healthcare data with a focus on scalability and performance. The position requires deep collaboration with cross-functional partners to translate business needs into technical data solutions. You will lead efforts to standardize data quality, observability, and monitoring practices across the platform. This role is central to enabling data-driven decision-making and product development across Komodo Health.
Key facts
What you'll do
- Architect and maintain large-scale production data pipelines using Python, SQL, and Airflow to support evolving business requirements.
- Transform complex healthcare datasets including claims, EHR, and reference data into structured, usable products for downstream consumers.
- Implement comprehensive data quality checks, lineage tracking, observability, and monitoring frameworks to ensure system integrity and reliability.
- Diagnose and resolve performance bottlenecks and intricate data issues within high-compute workflows to maintain efficient operations.
- Partner closely with Product, Engineering, and Platform teams to design scalable technical solutions that align with strategic objectives.
- Oversee CI/CD processes, technical documentation, and rotational production support to ensure smooth deployment and operation of data services.
- Leverage AI-augmented tools for coding assistance, debugging, documentation generation, and architectural analysis to accelerate development.
- Optimize data workflows for efficiency, reliability, and maintainability while adhering to best practices in software engineering and data management.
- Collaborate with data scientists and analysts to ensure that data products meet analytical needs and support accurate decision-making.
- Contribute to the development of standards and guidelines for data modeling, naming conventions, and pipeline architecture.
- Participate in on-call rotations to address production incidents and provide timely resolution to critical data issues.
- Explore and prototype new data technologies and techniques to enhance the Healthcare Map and support emerging use cases.
- Ensure that all data processing activities comply with internal policies and regulatory considerations related to healthcare data.
- Mentor junior engineers and contribute to a culture of continuous learning and improvement within the Data Foundations team.
Requirements
- Extensive experience working with healthcare data including claims, clinical, RWE, provider, or patient datasets in production environments.
- Demonstrated proficiency with coding systems such as ICD-10, CPT, NDC, or NPI and understanding of their practical applications.
- Hands-on experience building, debugging, and maintaining production-grade data pipelines at scale using modern data engineering tools.
- Advanced proficiency in Python and SQL with a strong track record of writing efficient, maintainable, and well-documented code.
- Solid experience with workflow orchestration tools such as Airflow to design, schedule, and monitor complex data workflows.
- Familiarity with Spark or similar distributed processing frameworks for handling large-scale data transformations.
- Proven ability to design, deploy, and operate data infrastructure within AWS environments, leveraging core services effectively.
- Strong aptitude for root-cause analysis and production troubleshooting with a methodical approach to resolving complex issues.
- Experience with version control systems, code review processes, and collaborative software development practices.
- Understanding of data security, privacy, and compliance considerations relevant to handling sensitive healthcare information.
- Excellent problem-solving skills and the ability to work independently and collaboratively in a fast-paced environment.
- Strong communication skills to articulate technical concepts to both technical and non-technical stakeholders.
- Willingness to adapt to changing priorities and requirements in a dynamic, high-growth technology company.
- Commitment to maintaining high standards of code quality, testing, and documentation throughout the development lifecycle.
- Ability to manage multiple priorities and deliverables simultaneously while maintaining attention to detail.
Nice to have
- Experience delivering external-facing data products via APIs or serving layers to enable integration with third-party applications.
- Ability to optimize architectures for cost efficiency, versioning, and high-volume productization to support scaling demands.
- Experience applying agentic workflows or AI technologies to data engineering processes and operational tasks.
- Success working in high-growth, ambiguous environments where ownership and initiative are essential for driving progress.
- Familiarity with modern data stack components including data lakes, warehouses, and real-time processing platforms.
- Exposure to healthcare analytics, population health, or value-based care models is a distinct advantage.
- Track record of contributing to open source projects or engaging with the broader data engineering community.
Practical notes
- This position is based in the United States and offers remote, hybrid, or in-office work arrangements depending on location.
- Compensation for this role varies by location, with salary ranges set at $196,000 to $230,000 USD for SF Bay Area and NYC locations, and $170,000 to $200,000 USD for all other US locations.
- Benefits coverage includes health, dental, and vision insurance along with 401(k) contributions and company match.
- Additional benefits consist of flexible time off, disability and life insurance, equity awards, and performance-based bonuses.
- AI integration is a core expectation for daily tasks including documentation, code generation, and workflow automation to enhance productivity and accuracy.
- Professional development opportunities are supported to encourage continuous growth and skill advancement in data engineering technologies.
- The role may require occasional travel within assigned regions for team meetings, planning sessions, or company events as determined by organizational needs.
- Employment eligibility requires authorization to work in the United States without sponsorship for this position at this time.