Senior Software Engineer, Bulk/Interactive Inference
Job description
About the role
You will architect and evolve the inference platform that powers real-time and batched predictions for the Waymo Driver, ensuring it meets stringent latency and throughput targets under production loads. In this role, you will own the design of scalable serving systems that process petabytes of sensor data and model outputs daily while collaborating with data scientists to translate research models into robust production services. You will partner with cross-functional product teams to define and deliver infrastructure that accelerates experimentation and improves the reliability of hosted models. This position requires deep engagement with large language and perception models to optimize their execution across heterogeneous hardware. You will lead initiatives that reduce operational costs by improving resource utilization and inference efficiency at scale. You will also be responsible for debugging complex performance regressions and ensuring that critical inference pipelines remain resilient and observable around the clock. This is a hybrid role where your work will directly shape how the Waymo Driver accesses and utilizes intelligence from the entire ML stack.
Key facts
What you'll do
Architect and maintain the bulk and interactive inference infrastructure that serves as the backbone for Waymo's model deployment lifecycle.
Scale distributed serving systems to handle exabyte-scale datasets and high-concurrency inference requests across multiple geographic regions.
Optimize model execution paths to improve throughput and reduce latency for large language and multimodal models hosted on Waymo infrastructure.
Integrate and deploy inference solutions for diverse use cases including distillation, eval, dataset generation, active learning, and auto-labeling pipelines.
Collaborate closely with modeling teams to adapt model architectures for efficient hosting without sacrificing predictive quality.
Implement observability and debugging tooling to monitor inference performance, resource utilization, and cost metrics in real time.
Partner with product and infrastructure teams to ensure that hosted models meet reliability, security, and compliance standards for production use.
Drive automation of deployment, scaling, and rollback mechanisms to support rapid iteration and safe experimentation by ML engineers.
Requirements
Bring a minimum of 5 years of professional software engineering experience with a strong track record of delivering complex systems.
Demonstrate expertise in programming with C++ to build high-performance components that power inference workloads.
Show experience in designing and operating highly scalable distributed systems that remain performant under load.
Have a proven ability to work with datasets that grow to the order of exabytes while maintaining efficient data processing patterns.
Be comfortable deploying and tuning model serving stacks in production environments with strict latency and availability requirements.
Understand the trade-offs involved in running large language and perception models simultaneously on shared hardware resources.
Possess strong problem-solving skills to diagnose performance bottlenecks and implement fixes that scale with system growth.
Communicate effectively with both technical and non-technical stakeholders to align on goals and prioritize infrastructure improvements.
Nice to have
Bring passion for building internal infrastructure and tools that empower other engineers to ship faster.
Have hands-on experience with model hosting and inference optimization for transformer-based and other advanced architectures.
Have worked with data-intensive systems where exabyte-scale datasets are common and require careful handling.
Practical notes
This is a full-time position based in Mountain View, California, USA.
Candidates must be eligible to work in the United States without sponsorship for this role.
No specific hours, travel, visa, or application deadline details are provided in the source material.