Senior Data Scientist
Job description
About the role
You will own the end to end lifecycle of machine learning models that power analytics and intelligence features for healthcare revenue cycle management. You will conduct deep exploratory data analysis on large-scale claims datasets to surface patterns and anomalies that inform product decisions. You will collaborate closely with engineering teams to integrate models and data products into production systems reliably and at scale. You will monitor model performance in production and design iterative improvements based on observed results and stakeholder feedback. You will communicate complex analytical findings clearly to both technical and non-technical audiences including product and customer-facing teams. You will contribute to data pipeline and warehouse development efforts when necessary to support model requirements and data quality. This role is ideal for someone who executes well independently, takes ownership of ambiguous problems, and is ready to expand the scope of work over time. You will play a direct role in shaping how AI and analytics drive operational outcomes for healthcare revenue teams.
Key facts
What you'll do
- Build and deploy machine learning models that power our core analytics and intelligence features across the revenue cycle.
- Conduct exploratory data analysis on large-scale claims datasets to identify patterns, anomalies, and actionable insights.
- Collaborate with engineering to integrate models and data products into production systems with robust and scalable architectures.
- Monitor model performance in production environments and iterate based on measured results and business impact.
- Communicate findings clearly to cross-functional stakeholders including product managers and customer-facing teams.
- Contribute to data pipeline and warehouse development as needed to support modeling efforts and data reliability.
- Partner with product teams to translate business requirements into analytical approaches and model objectives.
- Leverage statistical analysis and experimentation to validate hypotheses and drive decision making.
- Implement solutions using ML frameworks such as scikit-learn, XGBoost, or similar tools.
- Utilize cloud infrastructure, preferably on AWS, to deploy and maintain scalable data science workflows.
- Work with data pipeline and orchestration tools like sqlMesh, Temporal, or equivalent platforms.
- Ensure models are monitored for drift, performance, and reliability in production settings.
- Translate complex analytical results into clear narratives for stakeholders with varying technical backgrounds.
- Support the development and maintenance of MLOps tooling including experiment tracking and model monitoring.
Requirements
- Eight plus years of experience in a data science or applied machine learning role in a professional setting.
- Demonstrated track record of taking models from development to production in a real product environment with measurable outcomes.
- Strong proficiency in Python for data analysis, modeling, and scripting alongside advanced SQL skills for querying large datasets.
- Experience successfully deploying at least one model from development through production stages in a live system.
- Solid grounding in statistics, probability, and machine learning fundamentals that underpin modern modeling practices.
- Hands-on experience with machine learning frameworks such as scikit-learn, XGBoost, or similar libraries.
- Experience with cloud infrastructure, preferably Amazon Web Services, for building and operating data systems.
- Experience with data pipeline tooling and orchestration frameworks such as sqlMesh, Temporal, or equivalent tools.
- Ability to communicate analytical findings clearly to both technical and non-technical audiences with precision and empathy.
- Experience with MLOps tooling including experiment tracking systems like MLflow, orchestration tools such as Temporal or Prefect, and model monitoring practices.
- Comfort working in fast paced, ambiguous environments where priorities evolve and problem solving is iterative.
- Commitment to writing clean, maintainable code and documenting work for reproducibility and knowledge sharing.
- Willingness to partner cross functionally with engineering, product, and business stakeholders to align on goals and outcomes.
Nice to have
- Exposure to electronic health record data or revenue cycle management data in prior roles.
- Familiarity with anomaly detection methods tailored to claims or transactional datasets.
- Familiarity with recommendation systems or ranking approaches in applied settings.
Practical notes
The role is based in New York City at our office located at 3 World Trade Center with easy access to all trains and the PATH, and offers amazing views of the city. This is a full time position with standard working hours. Candidates must be eligible to work in the United States without sponsorship for this role at this time. No relocation support is available. The compensation details were last updated on the date of this job page and remain subject to change based on business needs.