Senior Machine Learning Engineer
Job description
About the role
The role collaborates closely with model developers to ensure AI behavior is safe, reliable, and helpful for patients and practitioners. Core responsibilities center on defining evaluation strategies and running systematic experiments to measure reasoning, factuality, and user experience.
Data roles turn raw information into decisions. Analysts query databases and build dashboards. Data scientists build models that predict outcomes. Data engineers build the pipelines that move and store data. All three work closely with business teams and need a mix of statistics, coding, and communication. Nearly every modern company runs on data teams, from startups to banks. A strong portfolio of past analyses matters more than degrees in many hiring decisions.
Key facts
What you'll do
Evaluation strategies for AI agentic systems are defined and owned, covering metrics, protocols, datasets, and tooling. These strategies are designed to ensure AI healthcare solutions meet rigorous standards for safety and reliability across diverse use cases.
The pipelines provide objective measurements that help guide model improvements and validate system behavior before deployment.
Experiments are structured to isolate variables and generate actionable data that informs model adjustments.
Insights from evaluation results are translated for model developers and research scientists to drive iterative improvements in AI capabilities. Findings are communicated clearly to support data-driven refinements to models and evaluation protocols.
These contributions help standardize evaluation approaches across teams and promote consistent best practices.
Requirements
An MSc or PhD in Computer Science, Machine Learning, Data Science, or a related field is required to understand complex evaluation frameworks. Candidates must demonstrate deep knowledge of machine learning concepts and research methodologies relevant to healthcare AI.
7+ years of hands-on experience working with large language models such as GPT, Claude, Llama, or BERT-like architectures is required. Experience must include working with models in production or near-production environments.
Proven experience evaluating agentic or reasoning systems, such as autonomous agents, tool-using LLMs, dialogue systems, or task-oriented assistants, is required. This experience should include designing evaluation scenarios and interpreting system-level behaviors.
A strong track record in experiment design, metric definition, and evaluation automation is required. Candidates must show they can design rigorous tests and implement systems that collect meaningful performance data.
Ability to bridge research and production is required to influence both modeling approaches and product decisions effectively. This ensures evaluation insights translate into practical improvements in AI healthcare applications.
Excellent communication skills, a collaborative mindset, and fluency in English are required to work within cross-functional teams. Clear communication enables effective collaboration with model developers, researchers, and product stakeholders.
Nice to have
Experience in the clinical or medical domain and sensitivity to ethical or regulatory challenges in healthcare AI is noted as valuable. This background helps ensure evaluation practices align with healthcare-specific requirements and constraints.
Practical notes
Start date is as soon as possible. Typical interview steps
Data interviews commonly include a SQL or coding exercise, a statistics question, and a case study. Candidates may be asked to design a metric, interpret an experiment, or build a small model. Some companies give a take-home analysis. Expect questions about past projects and the business impact of your work. Interviewers often evaluate how you communicate uncertainty and business impact, not only the math. Bringing a clean write-up of a past analysis to the interview is well received.
Good to know
General machine learning evaluation roles often work with experiment tracking tools and monitoring dashboards to measure model performance over time. These roles typically require comfort with Python-based tooling and data analysis workflows. Evaluation methods in healthcare AI emphasize safety, bias detection, and compliance considerations. Modern agentic systems testing involves tracing decision paths and validating tool-use correctness. Continuous evaluation cycles support rapid iteration while maintaining reliability standards.