Senior Software Engineer, Infrastructure Automation and Distributed Systems
Job description
Senior Software Engineer, Infrastructure Automation and Distributed Systems at NVIDIA
About the role
We are looking for experienced Systems and Software Engineers to develop and maintain dependable large-scale infrastructure services. In this role, you will be responsible for ensuring that our internal and external-facing EDA services, built on NVIDIA hardware, operate reliably. If you are innovative and self-driven and enjoy tackling challenges, we would like to hear from you.
Key facts
What you'll do
- Create, implement, and manage infrastructure services while overseeing the software lifecycle to align with our business objectives.
- Contribute to defining service level objectives and error budgets for our internal services as part of our observability framework.
- Reduce repetitive tasks through automation where it is beneficial to build and maintain such solutions.
- Engage in blameless incident prevention and response as part of an on-call rotation.
- Advise and collaborate with peer teams on best practices for systems design.
Requirements
- Bachelor's degree in Computer Science or a related technical discipline that involves programming (such as physics or mathematics) or equivalent experience.
- A minimum of 12 years of relevant industry experience.
- Proven ability to initiate projects, persuade others to collaborate, and work effectively on projects initiated by teammates.
- Proficiency in infrastructure automation and distributed systems design, specifically in developing tools for large-scale private or public cloud environments in production.
- Familiarity with one or more programming languages such as Python, Go, Perl, or Ruby.
- Extensive knowledge of Linux, Networking, Storage, or Containers.
Nice to have
- A systematic approach to problem-solving, excellent communication skills, and a strong sense of ownership and initiative. Experience in enhancing business outcomes using coding assistants, MCP servers, or AI agents.
- Experience in developing or working with bare metal as a service (BMaaS) systems.
- Background in creating or managing multi-cloud infrastructure services and operating private or public cloud systems using Kubernetes, OpenStack, Docker, or Slurm.
- Experience in teaching reliability practices (e.g., SRE) or general cloud systems best practices to peers or external organizations (e.g., CRE).
- Familiarity with the NVIDIA Collective Communication Library (NCCL).
Your base salary will be determined based on your experience, location, and the compensation of employees in similar roles. The salary range for this position is $224,000 - $356,500 for Level 5 and $272,000 - $431,250 for Level 6.
You will also be eligible for equity and benefits.
Applications for this position will be accepted until at least August 4, 2026.
This job posting is for an existing vacancy.
NVIDIA utilizes AI tools in its recruitment processes.
NVIDIA is dedicated to creating an inclusive workplace and is proud to be an equal opportunity employer. We value diversity among our current and future employees and do not discriminate based on race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status, or any other characteristic protected by law.