
Senior Product Manager, Compute Platform
Job description
About the role
Roblox is seeking a Senior Product Manager to lead the strategy and execution of its Compute Platform, the foundational infrastructure powering AI models that enable millions of users to create, play, and connect in immersive 3D experiences. This role will define the roadmap for next-generation AI infrastructure, including the rapidly expanding fleet of GPUs and AI accelerators deployed across Roblox's core data centers, edge locations, and public cloud environments. The successful candidate will own the products that transform raw GPU hardware into reliable, production-ready AI compute infrastructure, managing everything from driver and firmware systems to fleet-wide health monitoring and the abstractions that product teams depend on. This position sits at the intersection of infrastructure, AI, and product strategy, shaping how Roblox scales its most critical workloads.
Key facts
What you'll do
- Define and execute the strategic vision and product roadmap for Compute Platform, encompassing Managed Kubernetes (Roblox Kubernetes Service), Managed Compute Services, critical distributed systems, and the entire fleet of GPU and CPU infrastructure managed through unified Fleet APIs across on-premises and cloud environments
- Lead the transformation of Compute infrastructure to accommodate Roblox's most mission-critical workloads, including AI model training and inference, storage systems, and data analytics platforms, each presenting unique performance and reliability requirements
- Architect and scale GPU infrastructure capable of supporting both training and inference for frontier AI models, enabling breakthrough innovation while maintaining optimal cost efficiency and resource utilization
- Maintain unwavering focus on Compute Platform reliability for AI and other essential Roblox workloads, implementing systems and processes that reduce mean-time-to-detection and mean-time-to-recovery when failures occur
- Collaborate extensively with seven key platform teams to deeply understand their specific use cases, requirements, and pain points, then deliver Compute Platform primitives that empower them to build sophisticated solutions efficiently
- Serve as the primary cross-functional leader and communication hub, connecting Compute Platform users, engineering teams, AI researchers, product managers, finance stakeholders, and executive leadership to ensure alignment and rapid decision-making
- Design and implement abstractions and interfaces that product teams across Roblox can build upon, ensuring consistency, reliability, and ease of use while hiding underlying infrastructure complexity
- Drive continuous improvement in fleet-wide health monitoring and performance optimization, establishing metrics and dashboards that provide visibility into system behavior and enable proactive intervention
- Balance competing priorities and constraints across multiple stakeholder groups, making informed tradeoffs between resource utilization, latency requirements, cost considerations, and time-to-market pressures while never compromising on reliability
- Champion a builder mindset throughout the organization, actively prototyping new capabilities, iterating rapidly based on user feedback, and leveraging AI tools for ideation and value creation
- Establish governance frameworks and best practices for GPU and accelerator scheduling, including topology-aware placement strategies, preemption policies, and hardware affinity constraints
- Partner with finance and operations teams to optimize the total cost of ownership for compute infrastructure, identifying opportunities for efficiency gains and cost reduction without sacrificing performance
Requirements
- Seven or more years of product management experience with specific focus on compute infrastructure or distributed systems operating at significant scale
- Deep practical knowledge of Kubernetes internals and control plane components, including understanding of API Server performance bottlenecks, Etcd scaling limitations, Kubelet behavior patterns, and other critical architectural considerations
- Demonstrated experience productizing custom Kubernetes Operators, Controllers, and Custom Resource Definitions (CRDs) to extend platform capabilities beyond standard vanilla Kubernetes implementations
- Comprehensive familiarity with GPU and accelerator architecture along with the unique scheduling challenges they present, including topology-aware placement, preemption mechanisms, hardware affinity constraints, and resource readiness validation to prevent wasted compute cycles
- Working knowledge of cloud-native service networking concepts including microservices architectures, Container Network Interfaces (CNIs), and enterprise-grade service mesh implementations
- Proven track record building production-grade compute platforms where efficiency, reliability, and developer experience were fundamental design principles from inception
- Exceptional ability to balance the needs of multiple users and stakeholders, making optimal tradeoffs between utilization efficiency, latency requirements, cost constraints, and delivery timelines while maintaining system reliability
- Strong builder mindset with genuine passion for prototyping new solutions, evolving products through rapid iteration cycles, and leveraging AI technologies for ideation and delivering user value
- Excellent cross-functional leadership and communication skills, with ability to influence without direct authority and build consensus across diverse technical and non-technical audiences
- Strategic thinking combined with tactical execution capabilities, able to define long-term vision while delivering incremental value
Nice to have
- Hands-on experience building and operating compute infrastructure on major cloud platforms such as AWS, Google Cloud Platform (GCP), or Microsoft Azure
- Background in AI model development, training pipelines, or inference serving, providing firsthand understanding of the compute requirements and constraints AI workloads present
- Kernel-level programming experience or familiarity with custom kernel drivers, enabling deeper technical conversations with infrastructure engineering teams
- Experience designing or implementing agentic systems specifically for compute or infrastructure management use cases
- Prior work in high-growth technology companies or platforms serving developer communities
- Track record of managing large-scale infrastructure migrations or platform modernization initiatives
Skills & tools
Kubernetes, Kubernetes Operators, Custom Resource Definitions (CRDs), GPU architecture, AI accelerators, distributed systems, fleet management, API design, cloud platforms (AWS/GCP/Azure), service mesh, CNI, microservices, infrastructure reliability, cost optimization, cross-functional leadership, product strategy, roadmap planning, stakeholder management
Practical notes
This position is based at Roblox headquarters in San Mateo, California, with an expectation of onsite presence Tuesday, Wednesday, and Thursday each week. Monday and Friday office presence is optional unless otherwise specified. The compensation package includes a base salary between $280,540 and $330,950 USD annually, with the actual offer dependent on professional background, training, work experience, business needs, and market demand. All full-time employees receive equity compensation in addition to base salary and are eligible for comprehensive benefits. Roblox is committed to equal employment oppo