Solutions Architect, Infrastructure
NVIDIA AIUSA3w ago
SolutionsArchitectInfrastructureremotecurated-jd
Job description
Solutions Architect, Infrastructure at NVIDIA AI.
About the role
This position involves managing the deployment and integration of next-generation Data Center GPUs and networking platforms for major global customers. You will serve as the primary technical bridge between product strategy, cloud engineering, and large-scale infrastructure implementation.
Key facts
What you'll do
- Manage the end-to-end launch of Data Center GPU and networking platforms for Hyperscaler clients.
- Align with product teams to define success metrics and ensure technical consistency across sales and engineering departments.
- Coordinate across internal groups and external customers to remove obstacles and accelerate deployment timelines.
- Monitor deployment data to track product health, identify operational risks, and detect system bottlenecks.
- Troubleshoot complex issues involving firmware, drivers, containers, and distributed system interactions.
- Provide status updates and decision-making support to executive leadership.
- Contribute to future platform design and operational workflows through feedback to engineering teams.
Requirements
- BS, MS, or PhD in Computer Science, Electrical Engineering, Physics, or a related field, or equivalent practical experience.
- Minimum of 4 years in infrastructure engineering or solutions architecture.
- Practical experience validating and bringing up large-scale NVIDIA GPU platforms, including multi-node and multi-GPU setups.
- Knowledge of high-performance networking, including RDMA and congestion control for AI workloads.
- Proficiency with NVIDIA software stacks such as CUDA, NCCL, NVLink, and NVSwitch.
- Expertise in Linux performance tools like nvidia-smi, ethtool, perf, numactl, lspci, and container-level utilities.
- Deep understanding of server hardware, including PCIe topologies, BIOS/UEFI, NUMA, and thermal management.
- Experience with out-of-band management via Redfish, IPMI, or BMC.
- Strong grasp of Linux kernel subsystems, cgroups, and driver architecture.
Nice to have
- Experience with cloud infrastructure, including instance types and networking primitives at major Cloud Service Providers.
- Proven ability to lead cross-functional teams through complex infrastructure challenges.
- Background in transitioning hardware products from pilot phases to high-volume data center production.
- Familiarity with distributed training, inference, and modern LLM architectures.
Skills & tools
- Linux (drivers, kernel, cgroups, containers)
- NVIDIA stack (CUDA, NCCL, NVLink, NVSwitch)
- Hardware management (BMC, IPMI, Redfish, BIOS/UEFI)
- Networking (RDMA, high-bandwidth interconnects)
- Performance analysis (dmesg, journalctl, lspci, numactl, ethtool, iostat, perf, nvidia-smi, top/htop)
Practical notes
- Salary range for Level 3: 152,000 USD to 241,500 USD.
- Salary range for Level 4: 184,000 USD to 287,500 USD.
- Compensation includes equity and benefits.
- Application deadline: July 11, 2026.
- NVIDIA utilizes AI tools during the recruitment process.
- NVIDIA is an equal opportunity employer.