Staff Machine Learning Engineer
Job description
About the role
Staff Machine Learning Engineer
Waymo builds the technologies that enable fully autonomous driving. The company began as the Google Self-Driving Car Project in 2009. It created the Waymo Driver, described as The World's Most Experienced Driver. This system enables fully autonomous ride-hail service and can be applied to different vehicle platforms and product use cases. The Waymo Driver has completed over ten million rider-only trips. It operates using experience from more than 100 million miles driven on public roads and tens of billions of miles in simulation across 15 or more U.S. states.
The DUE Machine Learning team builds and operates scalable machine learning and data systems. It develops simulation workflow and insight tools. The team improves and accelerates evaluation and developer onboarding. It combines expert human judgment with advanced machine learning models. These models deliver training and evaluation data for hundreds of metrics and components that form the Waymo Driver. The team seeks researchers and software engineers passionate about machine learning for evaluation systems. The team requires an relentless drive to improve the performance of the technology stack.
This is a Staff Machine Learning Engineer position in Mountain View, California. The role is full-time. Base compensation is $183,900.00 per year.
What you will do
You will design ingestion pipelines that prepare sensor and simulation inputs for model evaluation. You will architect training workflows that enable large-scale model fine-tuning for driving behavior assessment. You will guide architectural choices for data and model layers spanning simulation, evaluation, and deployment. You will oversee production optimization of models that assess millions of vehicle miles driven daily.
You will design and scale large distributed systems covering the ML lifecycle. These systems will support planet-scale dataset generation, model training, and evaluation. You will collaborate cross-functionally to derive performance and system-level requirements for large ML systems. You will translate product and business goals into measurable technical deliverables. You will ensure system component alignment.
You will provide deep technical leadership on large-scale ML model architectures, specifically for autonomous vehicle models. You will build scalable systems for training and fine-tuning large-scale models to evaluate interesting driving behaviors. You will provide guidance on architectural decisions and technical directions. You will own large, complex systems and drive architectures that meet technical and business objectives. You will champion tooling for developer journeys. This work will accelerate how teams validate and iterate on models. You will champion metrics design that captures complex behavior for advanced driver systems.
Requirements
You hold an M.S. or Ph.D. degree in Computer Science, Machine Learning, Artificial Intelligence, or a related technical field. Equivalent practical experience is also acceptable. You bring seven or more years of professional software engineering experience. This must include at least three years with large ML systems. You have a record of contributions to machine learning tooling and frameworks. Examples include PyTorch, Jax, TensorFlow, or Ray. You understand both the user-facing API and the internal workings.
You excel at distributed training techniques. This includes gradient sharding and optimization strategies for scaling large models. You use ML accelerator profiling tools to uncover performance bottlenecks. You have a deep understanding of state-of-the-art machine learning models such as autoregressive transformers. You have strong leadership skills. You have experience navigating cross-functional teams and providing technical leadership across multiple organizations.
Preferred qualifications
We prefer ten or more years of professional software engineering experience. This must include at least five years in machine learning infrastructure. Experience in the autonomous vehicles domain, robotics, or complex simulation environments is preferred. You have a deep understanding of state-of-the-art reinforcement learning techniques. This includes methods used for fine-tuning large models, such as from human feedback or preferences.
You have familiarity with large-scale simulation platforms and their integration with ML training workflows. You have experience designing and using metrics for evaluating complex AI systems. You have a track record of technical leadership. You influence senior stakeholders and drive innovation across team boundaries. You communicate complex technical concepts clearly to varied audiences.
Skills and tools
You work with Waymo Driver, TensorFlow, PyTorch, Jax, and Ray.