
Machine Learning Engineer, Fleet Monitoring & Response
Job description
About the role
Waymo seeks a Machine Learning Engineer, Fleet Monitoring & Response to own the design, scaling, and optimization of the real-time monitoring engine that supports the Waymo Driver across its expanding global operating locations. You will develop and deploy spatial-temporal anomaly detection models, including S2-cell statistical regressions, and leverage Multimodal Foundation Models such as Gemini and VLMs to detect, triage, and automatically respond to off-nominal operations. This role involves building and standardizing the ML infrastructure that powers fleet monitoring models, including automated training and inference pipelines, low-latency spatial data stores like in-memory S2 grids, and continuous model drift monitoring. You will partner closely with Product Data Scientists to productionize, evaluate, and scale experimental models, translating notebooks and prototype algorithms into robust production-grade systems. The position is integral to improving access to mobility while advancing the safety and reliability of the Waymo Driver. You will contribute directly to maintaining Waymo's leadership in autonomous driving technology and its mission to save thousands of lives lost to traffic crashes. This role requires ownership of end-to-end system design and collaboration across engineering and product teams to ensure rapid response to fleet events.
Key facts
What you'll do
Design, scale, and optimize Waymo's real-time Fleet Monitoring and event response engine to support expansion to global operating locations.
Develop and deploy spatial-temporal anomaly detection models (e.g., S2-cell statistical regressions) and leverage Multimodal Foundation Models (Gemini/VLMs) to detect, triage, and automatically respond to off-nominal operations.
Build and standardize the ML infrastructure for fleet monitoring models, including automated training/inference pipelines, low-latency spatial data stores (e.g., in-memory S2 grids), and continuous model drift monitoring.
Partner with Product Data Scientists to productionize, evaluate, and scale experimental models, translating notebooks and prototype algorithms into production-grade systems.
Implement monitoring and observability frameworks to ensure high reliability and performance of the fleet monitoring stack in production environments.
Collaborate with cross-functional teams to define key operational metrics and drive improvements in detection accuracy and response time.
Develop data-driven insights from fleet telemetry to inform product decisions and enhance the overall rider experience.
Contribute to the architecture and evolution of the platform supporting large-scale simulation and real-world data integration.
Support the creation of automated response mechanisms that leverage multimodal inputs for rapid mitigation of edge-case scenarios.
Lead technical investigations into failures and near-misses using data from the fleet monitoring systems to inform model and system improvements.
Champion best practices in software engineering, machine learning operations, and data governance across the team.
Work within a fast-paced environment to prioritize initiatives based on impact to safety, operational efficiency, and product scalability.
Engage with external partners and internal stakeholders to align monitoring capabilities with evolving business and regulatory requirements.
Contribute to open-source tools and internal libraries where applicable to accelerate development and knowledge sharing.
Requirements
BS degree in Computer Science or equivalent practical experience.
5+ years of experience programming in backend coding languages such as Java or C++.
Experience in building backend platforms supporting multiple product use-cases/services.
Prior Machine Learning Engineering experience in Python using mature ML frameworks such as TensorFlow, PyTorch or Keras.
Strong understanding of software engineering principles, including version control, testing, and code review.
Proficiency in writing clean, maintainable, and scalable code to support long-term product needs.
Ability to work effectively in a distributed team environment across multiple time zones.
Willingness to relocate to or continue working in designated Waymo office locations as required by the role.
Demonstrated experience with data-intensive systems and real-time processing pipelines.
Commitment to following Waymo's safety-critical standards and operational practices.
Nice to have
MS in Computer Science, or equivalent practical experience.
Experience building and deploying ML / Optimization models into production environments.
Experience developing ML data pipelines and ML workflow automation code on top of a mature ML infra.
Experience working at another Ride hailing or Marketplace company.
Coursework background in ML and Optimization.
Practical notes
This is a full-time position based in designated office locations. Compensation details are provided for transparency and may vary based on location and individual qualifications. Relocation support may be available for eligible candidates. Candidates must be authorized to work in the country where the position is located.