Product Lead, AI/ML
Job description
About the role
Abridge is seeking a Product Lead, AI/ML to own the strategy and execution of our core measurement and evaluation infrastructure that underpins every product line. You will architect the platform that enables rapid, trustworthy experimentation with frontier models while preserving the scientific rigor clinicians depend on. This role owns the end-to-end eval lifecycle from data curation through production monitoring, establishing the standards that make AI quality observable and actionable. You will define the guardrails and governance that allow teams to move quickly without compromising safety, reliability, or trust. The hire will act as the central point of accountability for ensuring that evaluation practices remain model-agnostic, evidence-based, and aligned with clinical reality. You will translate complex measurement challenges into clear product requirements that balance ambitious research needs with the realities of deployed healthcare environments. This position is responsible for turning abstract quality goals into concrete platform capabilities that every team can use consistently.
Key facts
What you'll do
- Drive product strategy and execution for the evals platform, owning the roadmap across the eval lifecycle and delivering against measurable outcomes.
- Build the shared measurement infrastructure that lets any pod define quality, run experiments, compare models, and monitor production behavior in real time.
- Make model selection fast and routine by establishing a repeatable evaluation process for new frontier models within days of release.
- Own the operating model for evaluation gates from early build through GA and steady-state monitoring, clarifying which gates are hard versus advisory.
- Define and maintain standards for LLM judges, rule-based evaluators, human annotation, and online monitoring, and specify appropriate use cases for each.
- Work cross-functionally with Data Engineering on de-identification pipelines, Data Science on bootstrapping judge quality with reduced human annotation, and Clinical Science on flagged production cases.
- Collaborate with the agent platform team as workflows become more agentic, ensuring eval infrastructure keeps pace with autonomous system behaviors.
- Partner closely with ML researchers and engineers to drive impact in production, translating research advances into reliable evaluation methodologies.
- Balance long-term architectural investments with near-term quality improvements, ensuring the platform supports both innovation and stability.
- Operate with a high bar for quality, speed, and accountability, maintaining neutrality as the referent platform for model evaluation across the company.
- Define defensible frameworks for when a post-trained specialty model is worth it versus a prompted frontier model, based on measurable tradeoffs.
- Keep the eval system model-agnostic so it remains a neutral referee that works across architectures, modalities, and deployment strategies.
- Land evaluation standards across product pods you do not own, aligning on critical-error floors and non-negotiable quality thresholds.
- Enable teams to iterate quickly on models and metrics without breaking clinician trust or regulatory confidence.
- Maintain clear narratives about why specific evaluation approaches were chosen and how they map to real-world clinical outcomes.
Requirements
- 5+ years of product management experience with significant ownership of ML powered products or platform systems.
- Deep understanding of how to measure and improve model quality, including evaluation frameworks, annotation pipelines, and benchmark design.
- Strong technical fluency across ML, data pipelines, and distributed systems.
- Experience working closely with ML researchers and engineers to drive impact in production environments.
- Ability to balance long term architectural investments with near term quality improvements while maintaining rigorous standards.
- Strong communication skills and the ability to translate complex technical concepts into clear decisions and narratives for diverse stakeholders.
- A track record of delivering high quality products in domains where accuracy, reliability, and trust are paramount.
- Comfort working in ambiguous, fast-moving environments while upholding standards that protect patient safety and clinical validity.
Nice to have
- Experience building evaluation platforms, ML observability systems, or quality measurement pipelines.
- Work in clinical, healthcare, or regulated environments with a high bar for accuracy and compliance.
- Experience with specialty specific or domain specific model adaptations.
- Background in personalization systems, context ingestion frameworks, or ambient intelligence products.
- Experience shipping large-scale ML products with human-in-the-loop workflows and multi-stakeholder review processes.
Practical notes
Work is full_time based in the San Francisco office. The role requires close collaboration across teams and may involve occasional travel within the Bay Area for alignment and user research. No visa sponsorship is mentioned in the source text. There are no explicit compensation details provided in the source.