Principal Model Optimization Engineer
Job description
About the role
Roblox is seeking a Principal Model Optimization Engineer to join their ML Platform team, which currently powers hundreds of machine learning applications and processes billions of inference requests daily across discovery, safety, engine functionality, and numerous other domains. In this senior technical position, you will dive deep into the internal architecture of machine learning models to dramatically improve their performance during both training and inference phases. The role demands an engineer who thrives on extracting maximum efficiency from hardware and algorithms while building scalable solutions that benefit the entire organization. You will play a critical part in shaping how Roblox delivers immersive experiences to tens of millions of daily users who come to the platform to explore, create, play, learn, and connect with friends in three-dimensional digital environments.
Key facts
What you'll do
- Enhance machine learning model performance specifically targeting GPU hardware architectures, addressing both the training pipeline and real-time inference workloads to reduce computational overhead and improve throughput.
- Perform granular performance profiling and analysis at the system level to uncover bottlenecks within existing machine learning workflows, then develop concrete recommendations and implementations to resolve these issues.
- Build and maintain best practices documentation along with specialized tooling that enables consistent model optimization and streamlined deployment processes across the organization.
- Work closely with multidisciplinary teams including data scientists, machine learning engineers, and software developers to ensure optimized models integrate seamlessly into production systems serving millions of users.
- Partner with teams across different organizational units to create intuitive tooling, well-designed interfaces, and informative visualizations that make the machine learning platform at Roblox enjoyable and efficient to use.
- Apply advanced optimization techniques specifically tailored for large language models, including speculative decoding strategies, continuous batching implementations, and various quantization approaches to reduce model size while preserving accuracy.
- Investigate and resolve GPU-related issues by interpreting detailed GPU profiles, diagnosing Xid errors, and implementing fixes that ensure stable and efficient hardware utilization across the inference fleet.
- Develop and refine frameworks that deliver repeatable performance improvements, ensuring that optimization gains can be systematically applied across multiple models and use cases rather than requiring bespoke solutions each time.
- Contribute to reducing inference latency and improving model execution speed through expert application of specialized tools and frameworks designed for GPU acceleration.
- Support internal stakeholders by understanding their specific requirements and translating those needs into platform capabilities that accelerate their machine learning projects.
- Explore innovative techniques and emerging technologies that could further enhance model performance, staying current with the rapidly evolving landscape of machine learning optimization.
- Participate in architectural decisions that affect the broader ML Platform, drawing on extensive system design experience to ensure solutions scale effectively across all of Roblox.
Requirements
- Minimum of six years of professional engineering experience with a substantial portfolio of system design work demonstrating the ability to build high-performance systems at scale.
- Extensive hands-on experience debugging GPU-related issues, including the ability to read and interpret GPU performance profiles and diagnose hardware-level errors such as Xid faults.
- Strong proficiency with advanced acceleration tools and frameworks including CUDA for GPU programming, Triton for kernel development, and TensorRT for inference optimization.
- Demonstrated expertise applying optimization techniques to large language models, with practical experience implementing speculative decoding, continuous batching, quantization, and similar approaches.
- Deep passion for performance engineering with a track record of pushing systems to their theoretical limits, whether through GPU optimization, algorithmic improvements, or novel performance enhancement techniques.
- Strong commitment to building generalized tools and frameworks rather than one-off solutions, ensuring that performance improvements benefit multiple teams and use cases across the organization.
- Genuine enthusiasm for supporting internal partners such as data scientists and machine learning engineers, with excellent communication skills to understand and address their technical needs.
- Bachelor's degree in Computer Science, Computer Engineering, Data Science, or a closely related technical discipline.
Nice to have
- Experience working on machine learning platforms that serve billions of daily inference requests across diverse use cases and application domains.
- Background in building developer-focused tooling and interfaces that prioritize usability and developer experience alongside raw performance.
- Familiarity with the unique challenges of optimizing models for real-time applications in gaming or interactive entertainment contexts.
- Prior experience at companies operating large-scale consumer platforms with demanding latency and throughput requirements.
- Knowledge of emerging optimization techniques and hardware architectures that could provide future performance advantages.
- Experience mentoring other engineers on performance optimization best practices and GPU debugging methodologies.
Skills & tools
- Programming: CUDA, Python, C++
- Frameworks: Triton, TensorRT, PyTorch
- Optimization techniques: Speculative decoding, continuous batching, quantization, kernel optimization
- Profiling: GPU performance profiling tools, system-level debugging
- Infrastructure: Large-scale ML platform development, production deployment pipelines
- Collaboration: Cross-functional partnership, internal developer support, technical documentation
Practical notes
- The position is based at Roblox headquarters in San Mateo, California, with a hybrid schedule requiring in-office presence on Tuesday, Wednesday, and Thursday each week.
- Compensation includes a competitive base salary between $295,250 and $345,040 USD depending on factors such as professional background, training, work experience, location considerations, business requirements, and market conditions.
- Full-time employees receive equity compensation in addition to base salary, along with a comprehensive benefits package.
- Candidates should be aware that certain visa categories may present challenges for employment, and the company may not be positioned to provide H-1B sponsorship for this role at the current time.
- Roblox maintains a commitment to equal employment opportunity and provides reasonable accommodations for candidates with qualifying disabilities or religious observances throughout the recruiting process.
- The salary range disclosed is subject to modification and may be adjusted in the future based on market conditions and business needs.
About the company
An online space lets people build and play games. Founded long ago by two creators, it grew into a shared world opened to everyone in the mid 2000s. Now reaching millions each day, it hosts games made by its users. A large share of young American children under sixteen are monthly players here. High daily activity shows a lasting community where creativity and shared play remain central.