Principal Software Engineer, Profiling Services
NVIDIAUSA1w ago
Engineeringremotecurated-jd
Job description
Principal Software Engineer, Profiling Services at NVIDIA.
About the role
You will architect and build a production-grade, low-overhead GPU profiling service designed to operate continuously across large-scale clusters. This role focuses on delivering deep performance insights for machine learning workloads by integrating across system software, drivers, and the CUDA stack.
Key facts
What you'll do
- Define the architecture, data pipelines, and scalability standards for a profiling service that monitors multiple processes, GPUs, and nodes.
- Write high-reliability C/C++ code that manages shared memory and IPC while maintaining strict limits on CPU and memory consumption.
- Manage the full lifecycle of feature development from user-mode components down to driver-level performance counters and trace providers.
- Create models that connect low-level system data to ML frameworks like PyTorch and XLA to provide clear performance feedback.
- Guide the technical strategy for the engineering team, mentor staff, manage architectural risks, and coordinate roadmaps with cross-functional partners.
Requirements
- BS or MS degree in Computer Science, Computer Engineering, or a related field.
- 15+ years of experience in system-level C/C++ development, specifically regarding performance engineering, concurrency, and memory management.
- History of shipping production-ready system software or drivers that require high reliability and observability.
- Experience in technical leadership, including setting success metrics, defining system architecture, and managing roadmaps within fast-paced teams.
- Strong communication skills to influence stakeholders and build relationships across organizational boundaries.
Nice to have
- Background in CPU/GPU tracing stacks such as Nsight, CUPTI, event correlation, and performance counters.
- Deep knowledge of GPU architecture, CUDA streams, graphs, runtime APIs, and kernel behavior.
- Experience building always-on or multi-client profiling tools that maintain predictable overhead at scale.
- Familiarity with ML ecosystems like JAX or PyTorch and experience tuning training and inference loops by identifying compute or memory bottlenecks.
- Experience with user-mode driver development and platform security models.
Skills & tools
- C/C++
- CUDA
- PyTorch
- JAX
- XLA
- Nsight
- CUPTI
- IPC/Shared Memory
Practical notes
- Base salary range: 272,000 USD to 431,250 USD, depending on location and experience.
- Compensation includes equity and benefits.
- Application deadline: February 14, 2026.
- NVIDIA utilizes AI tools during the recruitment process.
- NVIDIA is an equal opportunity employer committed to a diverse workplace and does not discriminate based on protected characteristics.
- Job ID: JR2009909.