Senior Machine Learning Engineer, Reliability
Job description
About the role
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences - all created by our global community of developers and creators. At Roblox, we're building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We're on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you'll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. The Reliability team at Roblox operates at the depth and breadth of the Roblox stack. Availability of the platform is a key company goal. We are hiring our first Senior Machine Learning engineer within our team. As a Senior Machine Learning Engineer within Reliability, you will help set the direction for how machine learning systems/practices can be leveraged to improve the reliability of the overall Roblox platform. You will own the architectural and execution roadmap of leveraging massive data across - logs, traces, metrics, production changes, to proactively detect issues before they become real problems (MTTD) and/or reduce time to resolve incidents (MTTR). You will have the opportunity to cross functionally collaborate with other similar teams at Roblox to define best practices and software.
What you'll do
Help define the roadmap for leveraging Machine Learning Engineering to improve Production Systems Reliability at Roblox. Improve realtime anomaly detection capabilities by leveraging various state of the art ML techniques, thereby directly contributing to improving Mean Time to Detect Production issues. Develop methods to build pipelines to consume various streams of data (metrics, logs, traces, change management systems etc.). Build a reasoning layer that interacts with the streams of data to find possible root causes of problems happening in production. Build time-series models to predict capacity exhaustion and seasonal traffic spikes to drive automated scaling. Implement feedback loops from production signals to refine models continuously in a dynamic environment. Champion experimentation to validate the impact of novel ML approaches on reliability outcomes. Partner with SRE and platform teams to codify reliability policies into machine actionable signals. Design monitoring dashboards that translate model outputs into actionable reliability insights for operators. Contribute to open source reliability tooling where appropriate to amplify industry best practices. Mentor junior engineers on ML best practices and reliability oriented model development. Ensure model outputs are interpretable and actionable for incident responders.
Requirements
You have a BS, MS, or PhD in Computer Science, Statistics, or a related quantitative field, or equivalent practical experience. You have 5+ years of hands-on experience building ML models in production environments. You are proficient in Python and modern data science toolchains (e.g. pandas, scikit-learn, pyTorch, TensorFlow). You have experience with distributed systems and large scale data pipelines. You are comfortable working with logs, metrics, and traces in observability platforms. You have a strong grasp of statistical modeling and time-series analysis. You have experience with cloud platforms and infrastructure as code. You have excellent written and verbal communication skills.
Nice to have
Experience with anomaly detection frameworks and observability platforms. Knowledge of causal inference methods for production systems. Familiarity with Site Reliability Engineering practices and incident response processes. Contributions to open source projects in monitoring or reliability domains. Experience with container orchestration and service mesh technologies. Background in gaming, media, or high-scale consumer platforms.
Practical notes
This is a full-time position based in San Mateo, California, United States. Relocation assistance is not available for this role. The position requires authorization to work in the United States without sponsorship. Candidates must comply with all Roblox employment eligibility requirements. Interviews will be conducted onsite in San Mateo, CA.