Software Engineer, ML Infrastructure Platform
Job description
About the role
Nuro is seeking a Software Engineer to join the ML Infrastructure Platform team in a key role focused on the systems that power the Nuro Driver™. The successful candidate will own the infrastructure that trains the core models driving our autonomy stack, with direct impact on how quickly the company can iterate and deploy new capabilities. You will be responsible for designing and operating the distributed training systems, data pipelines, and orchestration workflows that turn raw autonomy data into improved models. This role sits at the intersection of machine learning and production engineering, where reliability and operational maturity are as important as scale. You will work closely with autonomy teams to ensure training pipelines are introspectable, reproducible, and efficient. When issues arise in training or data generation, you will own the investigation and resolution, directly influencing the pace of model improvement. Your contributions will ensure the fleet can train and deploy models safely and at scale across all Nuro use cases.
Key facts
What you'll do
Contribute to the design, operation, and scaling of Nuro's training infrastructure, spanning multi-generation accelerators and multi-cluster scheduling and orchestration.
Design and operate large-scale data pipelines, including batch and streaming ingestion, storage layout, and high-throughput data generation and storage for autonomy workloads.
Design and develop agentic-first ML workflows, creating data-to-training-to-evaluation pipelines that are introspectable, reproducible, and easy for autonomy teams to run and extend.
Own reliability for critical training and release pipelines by implementing robust instrumentation, meaningful alerting, and on-call and incident-response practices.
Implement and maintain the operational maturity of training systems, ensuring high GPU utilization and efficient cost management across the fleet.
Collaborate closely with autonomy engineers to understand workload requirements and translate them into infrastructure capabilities that accelerate model development.
Build and evolve the tooling for observability and debugging of ML workloads, enabling fast diagnosis of performance bottlenecks and training regressions.
Optimize data storage and access patterns to support high-throughput training workloads while controlling cost and complexity.
Partner with software engineers and researchers to create reusable frameworks and libraries that standardize best practices for ML infrastructure on the autonomy team.
Drive the adoption of Kubernetes-native orchestration for ML workloads, improving reliability and developer experience across the platform.
Lead initiatives to reduce infrastructure costs while improving reliability and scalability of the training systems.
Participate in on-call rotations to respond to training pipeline incidents and ensure rapid resolution with clear communication to stakeholders.
Contribute to the development of runbooks, operational documentation, and post-incident reviews to elevate the team's operational standards.
Engage with the broader ML and autonomy communities to identify new technologies and practices that can benefit Nuro's infrastructure.
Requirements
BS, MS, or PhD in Computer Science, Electrical Engineering, or a closely related field, plus 1+ years of relevant work experience.
Willingness to deep-dive into implementation details and to raise the technical and operational standards of the broader engineering organization.
A demonstrated ownership mindset, driving systems to operational maturity through monitoring, alerting, runbooks, and proactive improvements.
Strong proficiency in Python, with comfort in C++, Go, or a similar systems language for performance-critical components.
Hands-on experience running production infrastructure on Kubernetes, including cluster operations, networking, and storage.
Solid distributed-systems fundamentals, with the ability to reason about performance, failure modes, and reliability across a complex, multi-component system.
Experience with large-scale machine learning training workflows and the challenges of scaling compute-intensive workloads.
Familiarity with cloud platforms, particularly GCP, and infrastructure-as-code practices for managing distributed systems.
Nice to have
Strong working knowledge of GCP and prior experience in building large-scale data generation pipelines for autonomy.
Experience with Kubernetes-native orchestration for ML workloads, including job scheduling and resource management.
Depth in GPU / distributed training internals, including NCCL and collective communication patterns.
Familiarity with GPU and training observability tooling, including how to use these tools to diagnose real-world bottlenecks.
A track record of driving down infrastructure costs while simultaneously improving reliability and scalability.
Practical notes
This is a full-time position based in Mountain View, California. The role is eligible for an annual performance bonus, equity, and a competitive benefits package.