Software Engineer, ML Infrastructure
Job description
About the role
Nuro is a self-driving technology company founded in 2016 with a mission to make autonomy accessible to all. We are building the world's most scalable driver by combining cutting-edge AI with automotive-grade hardware. The Software Engineer, ML Infrastructure role is central to this mission. You will join a team dedicated to developing the core machine learning infrastructure that powers Nuro's autonomous driving technology. This position involves building and maintaining the systems essential for the complete lifecycle of ML models: from training and evaluation to deployment and monitoring. You will contribute directly to the foundational tools that enable our self-driving software to operate safely and reliably. The role requires close collaboration with cross-functional partners to translate high-level objectives into robust technical solutions. You will be responsible for the entire lifecycle of infrastructure components, from initial design through deployment and continuous iteration. The work involves solving challenging problems related to scale, performance, and reliability in a demanding production environment. This role is critical for ensuring that machine learning experiments can run efficiently and that results are reproducible. Your contributions will have a direct impact on the capabilities and safety of Nuro's autonomous vehicle platform.
Location: USA
Engagement: Full-time
What you'll do
In this role, you will design and implement scalable infrastructure for machine learning model development. A primary focus will be optimizing systems for data processing, model training, and inference to ensure high performance and reliability. You will build and maintain the core tools used for data ingestion, storage, and preprocessing, forming the backbone of our ML workflows. Developing monitoring and debugging capabilities will be essential to diagnose and resolve issues in production ML workflows swiftly. You will also implement automation for deploying machine learning models to various environments, streamlining the transition from development to production. A key responsibility includes working on the performance tuning of distributed training jobs to reduce iteration time and accelerate innovation. You will create interfaces that abstract complex underlying systems, providing simpler, more accessible tools for data scientists. Evaluating and recommending new technologies to improve the ML infrastructure stack is another critical function. You will partner with product teams to understand requirements and deliver tailored solutions that meet operational needs. Documenting system architectures and operational procedures is vital for ensuring maintainability and long-term success. Your work will also involve supporting the deployment of models that handle real-world driving scenarios safely and effectively. Where applicable, you will contribute to open source projects to advance the field of autonomous driving and share best practices.
- Designing and implementing scalable infrastructure for ML model development.
- Optimizing systems for data processing, model training, and inference.
- Building and maintaining core tools for data ingestion, storage, and preprocessing.
- Developing monitoring and debugging capabilities for production ML workflows.
- Implementing automation for model deployment across various environments.
- Performing performance tuning of distributed training jobs to reduce iteration time.
- Creating abstracted interfaces for simpler use by data scientists.
- Evaluating and recommending new technologies for the ML infrastructure stack.
- Partnering with product teams to understand requirements and deliver solutions.
- Documenting system architectures and operational procedures.
- Supporting the safe deployment of models for real-world driving scenarios.
- Contributing to open source projects to advance autonomous driving.
Requirements
To succeed in this role, you must have experience building and maintaining large-scale software systems. A demonstrated ability to work with complex data pipelines is essential, as you will manage critical data flows for ML operations. Strong problem-solving skills in a technical environment are required to navigate the challenges of scalable infrastructure. Proficiency in Python and C++ programming languages is mandatory, as these are the primary tools for development and optimization. A deep understanding of distributed systems principles and architectures is necessary to design robust and scalable solutions. Experience with cloud platforms such as AWS, GCP, or Azure is required to leverage modern infrastructure services effectively. You must have the ability to learn and adapt to new technologies quickly in a fast-evolving landscape. Strong written and verbal communication skills are crucial for effective teamwork and collaboration with cross-functional partners.
- Experience building and maintaining large-scale software systems.
- Demonstrated ability to work with complex data pipelines.
- Strong problem-solving skills in a technical environment.
- Proficiency in Python and C++ programming.
- Understanding of distributed systems principles and architectures.
- Experience with cloud platforms (e.g., AWS, GCP, Azure).
- Ability to learn and adapt to new technologies quickly.
- Strong written and verbal communication skills for effective teamwork.
Skills & tools
The successful candidate will possess expertise in the following core areas. Proficiency in Python is required for developing and scripting infrastructure components. Experience with C++ is essential for performance-critical system implementation. A solid understanding of distributed systems is necessary to build resilient and scalable architectures. Familiarity with major cloud platforms such as AWS, GCP, or Azure is required for deploying and managing infrastructure in production environments. These tools and skills form the foundation for enabling Nuro's mission to create a scalable driver through cutting-edge AI and automotive-grade hardware.