Senior Software Engineer, ML Infrastructure Platform
Job description
About the role
Nuro is a physical AI company pioneering Level 4 autonomy, and the ML Infrastructure role is central to that mission. The hire will own the core systems that train the models powering the Nuro Driver™, ensuring they are reliable, scalable, and efficient. This involves deep responsibility for distributed GPU training, data pipelines, and the orchestration that keeps the fleet running. You will directly impact the speed and quality of autonomy development by owning the critical path infrastructure. The role requires a strong focus on operational maturity, monitoring, and cost management. You will work closely with autonomy teams to build workflows that are introspectable and reproducible. This is a hands-on position where your systems will be the foundation for how the company ships self-driving capabilities.
Key facts
What you'll do
Design and operate the large-scale training infrastructure that spans multi-generation accelerators and multi-cluster scheduling and orchestration.
Architect and build robust data pipelines for both batch and streaming ingestion, defining storage layout and high-throughput data generation and storage strategies.
Develop agentic-first ML workflows, creating data-to-training-to-evaluation pipelines that are introspectable, reproducible, and easy for autonomy teams to run and extend.
Own the end-to-end reliability of critical training and release pipelines by implementing comprehensive instrumentation, meaningful alerting, and resilient on-call practices.
Drive down infrastructure costs while improving system reliability, establishing clear ownership for performance and operational efficiency.
Implement observability frameworks specifically for ML workloads, using data to diagnose real-time bottlenecks in GPU utilization and training throughput.
Collaborate closely with autonomy engineering teams to ensure the ML infrastructure evolves in direct support of the Nuro Driver™ development roadmap.
Define and maintain runbooks and operational standards that raise the maturity of the engineering organization.
Evaluate and integrate new technologies for distributed training, including deep expertise in NCCL and collective communication patterns.
Champion the use of Kubernetes-native orchestration for ML workloads, scaling the platform to meet the demands of the global mobility ecosystem.
Lead incident response for infrastructure issues, creating post-mortems and action plans that prevent regressions.
Partner with data science and model teams to align infrastructure capabilities with the needs of the next generation of autonomy models.
Requirements
BS, MS, or PhD in Computer Science, Electrical Engineering, or a closely related field, plus 3+ years of relevant work experience.
Willingness to deep-dive into implementation details and to raise the technical and operational standards of the broader engineering organization.
A demonstrated ownership mindset, driving systems to operational maturity through monitoring, alerting, runbooks, and clear on-call strategies.
Strong proficiency in Python, with comfort in C++, Go, or a similar systems language for performance-critical components.
Hands-on experience running production infrastructure on Kubernetes, including cluster operations and networking.
Solid distributed-systems fundamentals, with the ability to reason about performance, failure modes, and reliability across a complex, multi-cluster environment.
Experience with building large-scale data generation pipelines that handle high-throughput ingestion and storage layout optimization.
Deep understanding of GPU and distributed training internals, including collective communication, NCCL, and scaling strategies across multiple nodes.
Nice to have
Strong working knowledge of GCP and its core services for building and operating scalable infrastructure.
Experience with Kubernetes-native orchestration for ML workloads, including job scheduling and resource management.
Familiarity with GPU and training observability tooling, using dashboards and metrics to diagnose real bottlenecks.
A track record of driving down infrastructure cost while simultaneously improving reliability and system performance.
Practical notes
This is a full-time position based in Mountain View, California.
Eligible candidates must be authorized to work in the United States without sponsorship for this position.
No relocation assistance is provided.
The base pay range for this role is between $193,930 and $291,150, determined by factors including experience, qualifications, education, location, and skills. This position is also eligible for an annual performance bonus, equity, and a competitive benefits package.