Engineering Manager, Machine Learning Infrastructure, Ads
Job description
About the role
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences- all created by our global community of developers and creators. At Roblox, we're building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We're on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there. A career at Roblox means you'll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone. With Roblox Ads business growing at a rapid rate, we are building large scale ads machine learning infrastructure to deliver effective performance ads to our users, and more business values to our advertisers. We're looking for an EM to lead a team of exceptional ML infrastructure engineers, build scalable, reliable, and high-performance infrastructure that powers ML systems across our organization. You'll operate at the scales of hundreds of billions of engagements, and redefine how we deliver performance ads to hundreds of millions of users. You will own the end-to-end delivery of performant machine learning systems that power our advertising products, translating product goals into scalable technical solutions. You will establish robust engineering practices that ensure infrastructure reliability, efficiency, and cost optimization across the ML lifecycle. You will partner deeply with data scientists and product teams to remove roadblocks and accelerate high-impact experiments. You will foster a culture of ownership, continuous improvement, and technical excellence within your growing team. Your decisions will directly influence the reliability and scalability of the ads platform serving millions of users daily. You will play a key role in defining the technical vision and execution strategy for our ML infrastructure roadmap.
Key facts
What you'll do
Lead strategic planning and execution of scalable production-ready ML systems including model training, data pipelines, feature engineering and model inference.
Own the architecture, establish engineering best practices of scalability, reliability, and cost-effectiveness of ML infrastructure (e.g., training, serving, feature).
Work closely with data scientists, ML engineers, platform teams, and product stakeholders to design, implement, and operate robust ML platforms that accelerate model development and deployment.
Recruit, mentor, and grow a high-performing team of ML infrastructure engineers.
Define and drive technical standards for ML infrastructure to ensure consistency, reliability, and efficiency across the organization.
Collaborate with SRE and platform teams to ensure infrastructure reliability, performance, and observability at scale.
Evaluate, prototype, and recommend new technologies, tools, and frameworks to advance our ML infrastructure capabilities.
Partner with product managers to align infrastructure strategy with business objectives and product roadmaps.
Lead postmortem analysis and troubleshooting of complex infrastructure issues to drive continuous improvement.
Champion security, privacy, and compliance best practices within the ML infrastructure domain.
Mentor engineers through code reviews, design discussions, and career development to build a strong, scalable team.
Translate product requirements and data scientist experiments into robust, scalable ML infrastructure solutions.
Monitor and optimize infrastructure costs while maintaining high performance and availability.
Serve as a technical leader and representative for ML infrastructure across the organization and in cross-functional forums.
Drive the adoption of industry best practices for ML infrastructure, including MLOps, data versioning, and model lifecycle management.
Requirements
5+ years of experience designing, building, and deploying large-scale machine learning systems in production environments.
2+ years of experience managing ML infrastructure engineers.
End-to-End Execution: You've built ML systems from data ingestion to serving in production.
Team-first Mindset: A collaborative leader who thrives on mentoring and enabling others to succeed.
BS, MS, or Ph.D. in Computer Science, Engineering, or equivalent experience.
Experience with large-scale distributed systems and infrastructure.
Strong proficiency in at least one systems programming language such as Golang, C++, or Rust.
Deep understanding of machine learning workflows, including training, inference, and feature stores.
Proven ability to operate and troubleshoot complex systems in production environments.
Excellent communication and stakeholder management skills.
Nice to have
Experience with cloud platforms and infrastructure automation tools.
Knowledge of advertising technology systems and performance marketing concepts.
Experience with open-source ML infrastructure projects such as Kubeflow, MLflow, or TensorFlow Extended.
Practical notes
Roles that are based in an office are onsite Tuesday, Wednesday, and Thursday, with optional presence on Monday and Friday (unless otherwise noted).
Roblox provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws. Roblox also provides relocation support for candidates who are relocating for this role.