Staff ML Engineer, Perception Research
Job description
About the role
You design perception learning systems and own their full lifecycle from raw sensor input to deployed decision models. Your work shapes how the Waymo Driver sees and reacts across many cities and conditions. You translate research ideas into production code that runs on vehicle hardware at scale. You own the design, experimentation, and deployment of core perception learning systems end to end. You ensure that sensor data is transformed into reliable world models that drive safe and scalable autonomy. You work at the intersection of research innovation and production constraints to deliver robust perception capabilities. You collaborate closely with engineers and researchers to turn novel ideas into dependable software shipped in the Waymo Driver.
Key facts
What you'll do
Design scalable data pipelines that ingest and prepare training and evaluation inputs from camera, LiDAR, and radar.
Implement model distillation and bulk-inference infrastructure to move teacher knowledge efficiently into production services.
Conduct comprehensive experimentation to train and deploy multimodal LLMs and world models for 3D perception using sensor information.
Build and maintain tools for performance analysis, profiling (for example xprof), and debugging of ML models in demanding settings.
Develop and maintain evaluation frameworks that measure perception accuracy, safety, and robustness across diverse scenarios.
Experiment with model partitioning and sharding strategies to improve scalability and efficiency of globally distributed inference.
Partner with engineering and research teams across Alphabet to deploy new models and establish efficient workflows for continuous training.
Apply and develop techniques such as quantization, pruning, knowledge distillation, and efficient attention mechanisms for autonomous driving models.
Own the end to end lifecycle of perception models from raw sensor streams to deployed decision making components.
Champion best practices for data curation, labeling alignment, and sensor fusion within the perception training stack.
Drive innovation in representation learning that enables better generalization across cities, weather conditions, and edge cases.
Lead the design of infrastructure that supports efficient multimodal training and inference at vehicle scale.
Define and refine success metrics that connect perception model outputs to real world driving safety and performance.
Act as a technical leader in setting standards for model quality, reproducibility, and operational reliability.
Requirements
You hold a PhD or Masters in Computer Science, Machine Learning, Robotics, or a similar technical field, with four or more years of relevant experience.
You are proficient in implementing model training flows in a scalable, distributed, and performant manner, such as data parallel, FSDP, and other sharding approaches.
You demonstrate proficiency in JAX, Flax, and potentially TensorFlow or PyTorch for building and training models.
You are willing to work with the complexity of globally distributed inference infrastructure and production constraints.
You have hands on experience with optimizing the training and inference of Transformer architectures for real world systems.
You are comfortable designing systems that must meet strict latency, reliability, and safety requirements in production.
You can communicate clearly with both technical specialists and cross functional stakeholders about complex problems.
You are self directed and able to manage ownership of difficult technical problems with minimal supervision.
Nice to have
PhD in Computer Science, Machine Learning, or Robotics, with research focused on reinforcement learning, foundation models, or multi-modal learning.
Substantial involvement in and contributions to high impact industry AI projects that affect real vehicles.
Experience in generative models for domains such as world models, images, videos, and 3D using techniques like diffusion or autoregressive approaches.
Experience contributing to frameworks and libraries that improve training speed and scalability, for example JAX, Gemax, or XManager.
Skills & tools
Waymo Driver, TensorFlow, PyTorch, JAX, Flax, Xprof.
About the company
Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver-The World's Most Experienced Driver™-to improve access to mobility while saving thousands of lives now lost to traffic crashes.