Principal Software Engineer, GPU Compute
Job description
About the role
In this pivotal position at Roblox, you will serve as a key technical leader within the Compute team, focusing on GPU and AI accelerator capabilities. Your expertise will be crucial in managing the lifecycle of GPU hosts, ensuring they are production-ready and reliable for various workloads. This role involves tackling complex challenges that arise at scale, including driver management, GPU health, and performance optimization across a rapidly expanding fleet of accelerators. You will also play a significant role in shaping the technical direction of GPU compute within the organization.
Key facts
What you'll do
- Act as the primary technical authority on GPU technologies for the Compute team, collaborating with various teams including Kubernetes, Machine Bootstrap, Networking, and Cloud to develop a comprehensive GPU strategy.
- Oversee the lifecycle management of GPU hosts, which includes handling driver updates, firmware management, and ensuring GPU health and telemetry are monitored effectively.
- Design the architecture for how GPU resources are allocated to compute platforms, focusing on scheduling, resource isolation, and integration with Kubernetes for both GPU and AI workloads.
- Enhance GPU reliability and performance across the fleet by establishing protocols for the detection, diagnosis, and automated repair of malfunctioning accelerators before they disrupt production.
- Assess and integrate new GPU and AI accelerator platforms, as well as networking configurations such as NVLink, InfiniBand, and RoCE, to support multi-node training and inference.
- Create standards, tools, and APIs that enable other engineering teams to utilize GPU compute resources efficiently and safely, thereby minimizing manual effort and elevating overall performance.
- Lead initiatives to improve the organization's GPU expertise, mentoring engineers and fostering a culture of knowledge sharing and continuous improvement.
- Collaborate with cross-functional teams to ensure that GPU compute capabilities align with the broader objectives of the company and meet the needs of developers and creators.
Requirements
- A minimum of 10 years of experience in building and managing large-scale distributed systems and infrastructure.
- Extensive hands-on experience with GPU technologies, particularly in areas such as host provisioning, driver and firmware management, and ensuring GPU reliability in production environments.
- Proven track record in scaling GPU or accelerator infrastructure that supports critical workloads, demonstrating your expertise beyond mere fleet management.
- Strong programming skills in Go or other structured programming languages, with a focus on developing robust and efficient code.
- Experience with operating GPU and AI workloads in a production setting, including familiarity with CUDA, GPU scheduling, and high-performance networking technologies.
- Knowledge of Kubernetes for managing GPU workloads and an understanding of bare-metal concepts, including firmware and OS imaging, is highly desirable.
- A history of being recognized as the go-to expert for complex GPU and compute challenges, with the ability to lead and elevate the skills of your peers.
Nice to have
- Experience with advanced GPU architectures and emerging AI accelerator technologies.
- Familiarity with cloud-based GPU services and their integration into existing workflows.
- Previous involvement in open-source projects related to GPU computing or distributed systems.
Skills & tools
- Proficient in GPU management tools and frameworks, including CUDA and relevant monitoring software.
- Familiar with container orchestration platforms, particularly Kubernetes, and their application in GPU workloads.
- Strong understanding of networking technologies relevant to high-performance computing, such as NVLink and InfiniBand.
Practical notes
- The starting salary for this position ranges from $345,040 to $399,420 USD, depending on various factors including experience and market demand.
- This role is based at Roblox's headquarters in San Mateo, CA, with a hybrid work model requiring onsite presence three days a week, and optional attendance on Mondays and Fridays.
- Roblox is committed to providing equal employment opportunities and prohibits discrimination based on various characteristics. Reasonable accommodations are available for candidates with disabilities or religious beliefs during the hiring process.
- Please note that for U.S.-based roles, the company may have limitations regarding work authorization and may not support certain visa categories for this position.
Join Roblox and be part of a team that is redefining the way people connect and interact in immersive digital experiences. Your expertise in GPU compute will be instrumental in shaping the future of our platform and empowering a global community of creators.
About the company
An online space lets people build and play games. Founded long ago by two creators, it grew into a shared world opened to everyone in the mid 2000s. Now reaching millions each day, it hosts games made by its users. A large share of young American children under sixteen are monthly players here. High daily activity shows a lasting community where creativity and shared play remain central.