Staff Software Engineer, Data Products
Job description
About the role
Omada Health is on the lookout for a Staff Software Engineer to enhance and sustain the foundational data infrastructure that supports our machine learning efforts. In this position, you will be responsible for designing and implementing scalable data products and feature pipelines, which will empower data scientists and engineers to create and deploy intelligent solutions. This role plays a vital part in establishing dependable and reusable data foundations that facilitate personalized experiences for members and promote data-driven decision-making across the organization.
Key facts
What you'll do
- Design, develop, and manage reusable datasets that cater to various machine learning applications, including personalization, risk assessment, and recommendation systems.
- Create self-service tools and frameworks that simplify and expand access to dataset creation for the entire data organization.
- Work closely with Data Scientists to convert their modeling needs into production-ready feature pipelines, overseeing the entire model lifecycle from initial exploration to deployment.
- Identify essential source data, define necessary transformations, and establish historical data windows for effective feature engineering.
- Construct both batch and streaming data pipelines to convert raw healthcare, behavioral, product, and operational data into reliable datasets suitable for machine learning applications.
- Optimize large-scale distributed data processing to ensure efficiency, scalability, and cost-effectiveness.
- Maintain data quality through thorough testing, anomaly detection, schema validation, and pipeline monitoring.
- Lead architectural and design discussions for extensive machine learning data systems, advocating for standardized patterns and platform capabilities.
- Shape technical strategies across multiple engineering teams and collaborate with stakeholders from product, engineering, and business sectors to define data capture requirements.
- Mentor fellow engineers on distributed data processing, best practices in software engineering, and scalable data modeling techniques.
Requirements
- At least 8 years of experience in developing large-scale production data platforms and distributed data pipelines.
- Proven track record in designing reusable datasets that support machine learning, experimentation, or advanced analytics.
- Demonstrated capability to work closely with Data Scientists to implement feature engineering workflows in production.
- Experience leading technical initiatives across teams and influencing engineering direction effectively.
- Strong background in working with cloud-native data platforms, particularly AWS.
- Experience in building production data systems using technologies such as Databricks, Iceberg, Spark, Redshift, or Snowflake.
- Proficient in developing reliable batch and streaming data pipelines.
- Expert-level proficiency in SQL with robust data modeling skills.
- Strong programming capabilities in Python, Java, or Scala.
- Familiarity with Apache Spark or similar distributed computing frameworks.
- Experience with orchestration platforms like Airflow.
- Knowledge of feature stores or feature management platforms is a plus.
- A Bachelor's degree in Computer Science or a related field is preferred.
Nice to have
- Experience in the healthcare sector, particularly with behavioral or large-scale event data.
- Background in supporting personalization, recommendation, ranking, or predictive modeling systems.
- Familiarity with model training pipelines and MLOps workflows.
- Experience in designing data platforms for experimentation purposes.
- Knowledge of Data/AI Governance practices.
Skills & tools
Proficient in Python, SQL, Spark, Databricks, Iceberg, Redshift, Snowflake, Airflow, AWS, Kafka, Java, Scala, Docker, Kubernetes, Ruby on Rails, Postgres, Athena, Appflow, S3, SNS, SQS, Lambda, Serverless, Tableau, Bugsnag, Datadog, GitLabCI, Cursor, OpenMetadata, and NoSQL databases (including document and graph databases such as Neptune and Neo4j).
Practical notes
We offer a competitive salary complemented by a generous annual cash bonus. Employees can benefit from equity grants and an Employee Stock Purchase Plan (ESPP). Our remote-first culture allows for flexible work-from-home arrangements. We provide Flexible Time Off, generous parental leave, and comprehensive health, dental, and vision insurance. Additionally, we offer a 401k retirement savings plan, a Lifestyle Spending Account (LSA), and mental health support. Please note that visa sponsorship is not available for this position, and travel is not anticipated. The salary range for this role is between $210,000 and $260,000, depending on geographic location.