Senior Machine Learning Engineer, AI Platform
Job description
About the role
You will own the design and operation of the core AI infrastructure that powers WHOOP Coach and AI-powered Support, translating continuous physiological data into clear, actionable guidance for members. You will architect and maintain the evaluation pipelines, fine-tuning workflows, and LLM observability systems that ensure our models perform reliably in production. This role requires deep collaboration with data science and product teams to turn real member needs into robust, scalable AI systems. You will build the data curation and reshaping pipelines that convert messy, multi-source signals into high-quality training and evaluation datasets. In addition, you will establish feedback loops that use offline evaluations and real-world interactions to drive continuous model improvement. You will also mentor engineers and data scientists, elevating applied AI best practices across the AI Platform team. Finally, you will define and maintain the tooling that makes experimentation, deployment, and iteration faster, safer, and more repeatable.
Key facts
What you'll do
- Design, build, and operate production AI systems and scaffolding around language models that power conversational, predictive, and generative capabilities across WHOOP products.
- Lead end-to-end AI system initiatives spanning problem definition, data flows, dataset design, evaluation harnesses, deployment, and iteration in close partnership with data science and product.
- Build and maintain pipelines for collecting, curating, and reshaping messy, multi-source data into high-quality, well-structured training and evaluation datasets for language model-based systems.
- Operationalize fine-tuning and evaluation workflows for large language models behind member-facing features such as WHOOP Coach and AI Support, including defining datasets, labels, and taxonomies that reflect real member needs.
- Develop tooling and frameworks that make experimentation, offline/online evaluation, and model deployment faster, safer, and more repeatable, including robust observability for AI features in production.
- Build and maintain feedback loops that connect real member interactions, offline evaluations, and training data updates so that models improve continuously based on real-world behavior.
- Mentor other engineers and data scientists, share best practices in applied AI/ML, and help elevate the overall technical bar of the AI Platform team.
- Partner with cross-functional stakeholders to align AI capabilities with business goals, ensuring solutions are scalable, maintainable, and aligned with member value.
- Implement monitoring and alerting for AI systems to detect regressions, bias, and performance drift in production deployments.
- Contribute to the long-term architecture of the AI Studio at https://engineering.prod.whoop.com/ai-studio, ensuring components remain modular, testable, and secure.
- Define and track key quality metrics for language model outputs, aligning them with downstream product experiences and member outcomes.
- Collaborate closely with privacy and security teams to ensure compliance with data handling policies and best practices for sensitive information.
- Drive technical documentation for data pipelines, model interfaces, and evaluation frameworks to support knowledge sharing and onboarding.
- Explore and prototype new techniques for improving model efficiency, robustness, and personalization without compromising reliability.
- Participate in on-call rotations for AI platform services, supporting incident response and rapid resolution of production issues.
- Translate ambiguous product requirements into well-defined experiments, datasets, and evaluation protocols that de-risk AI releases.
- Establish benchmarks for model performance across different member segments, enabling data-driven decisions for future feature development.
- Enable other teams by providing reusable ML components, templates, and tooling that accelerate responsible AI delivery.
Requirements
- 3+ years of experience in applied machine learning, AI engineering, or ML-focused software engineering roles, including significant work in production environments.
- Hands-on experience building with modern language models (open-weight or API-based), including prompt design, fine-tuning, and rigorous evaluation.
- Solid working understanding of ML fundamentals (dataset construction, feature engineering, training workflows, evaluation metrics, experiment design) sufficient to make good engineering tradeoffs and partner effectively with data scientists.
- Familiarity with modern LLM training and alignment techniques such as supervised fine-tuning (SFT), direct preference optimization (DPO), and reinforcement learning (RL), and how they influence data requirements, evaluation strategies, and system design in production.
- Proven track record building, shipping, and operating ML-powered systems end to end, from data pipelines (batch and/or streaming) that transform large datasets into usable training and evaluation sets to production deployments with inference optimization, observability, and lifecycle management.
- Strong proficiency in data manipulation and analysis, including working with messy, multi-source, and semi-structured data and translating product questions into well-defined datasets, labels, and evaluation splits.
- Familiarity with best practices for secure, privacy-aware AI and working with sensitive data.
- Excellent communication and collaboration skills, with the ability to influence across teams and drive alignment on technical direction.
- Experience with version control, code review, and CI/CD practices for machine learning artifacts.
- Comfortability with cloud-native infrastructure, containerization, and orchestration tools commonly used in production ML systems.
- Ability to work independently and make sound technical decisions with ambiguous information while communicating tradeoffs clearly.
- Commitment to writing clean, maintainable, and tested code that can scale to support millions of members.
- Willingness to engage in continuous learning and adapt to rapidly evolving AI tools, techniques, and best practices.
Nice to have
- Preferred experience with real-time inference systems and performance optimization for LLM serving.
- Knowledge of MLOps platforms, experiment tracking, and model registry tooling.
- Experience contributing to open-source ML projects or publishing technical work in relevant domains.
- Background in health, wearables, or personalized continuous sensing data applications.
Practical notes
This role is based in the WHOOP office located in Boston, MA. The successful candidate must be prepared to relocate if necessary to work out of the Boston, MA office.