Engineering Manager, Ads ML Efficiency
Job description
Engineering Manager, Ads ML Efficiency
About the role
Reddit is establishing a specialized Ads ML Efficiency team to significantly improve the speed, cost-effectiveness, safety, and scalability of model training and inference. This role involves leading a group dedicated to optimizing models, enhancing training efficiency, enabling GPU usage, conducting load testing, developing model performance tools, and implementing efficiency safeguards across Ads ML.
Key facts
What you'll do
Lead and develop a team of ML and systems engineers focused on model optimization and ML efficiency.
Define the strategic direction for training and inference optimization, tooling for launch readiness, and reusable efficiency components for Ads ML.
Achieve measurable improvements in model training duration, online latency, serving expenses, and infrastructure-related launch risks.
Oversee the creation of systems for profiling, benchmarking, load testing, observability, cost analysis, debugging, and efficiency certification.
Collaborate with model owners and platform teams to expedite critical launches and remove production roadblocks.
Balance immediate optimization tasks with the long-term development of platform solutions and automation.
Foster alignment with cross-functional teams including MLP, AMP, Ranking, and serving to clarify responsibilities and advance Ads needs.
Promote engineering best practices in measurement, performance debugging, launch safety, and technical decision-making for efficiency initiatives.
Requirements
Extensive experience in ML engineering, with a deep understanding of model training, serving, debugging, and optimization.
Proven background in hands-on optimization, including improving training pipelines, serving systems, profiling processes, model/inference efficiency, or GPU utilization.
Demonstrated ability to build and lead teams, mentor engineers, manage project delivery, and navigate ambiguity in prioritization.
Proficiency in distributed systems, with experience reasoning about large-scale ML systems and their trade-offs regarding reliability, speed, cost, and scale.
A strong sense of customer focus and platform development, capable of serving modeling teams while building reusable systems.
Excellent communication skills, with the ability to articulate technical trade-offs to engineers, product managers, and senior leaders.
Experience in ads ranking, recommender systems, marketplace ML, or similar production ML fields is highly desirable.
Nice to have
Experience with GPU training and serving migrations.
Familiarity with PyTorch, distributed training frameworks, or low-level performance optimization.
Experience developing frameworks for efficiency benchmarking or launch certification.
Experience working in organizations with a separation between ML platform and applied modeling responsibilities.
Skills & tools
ML optimization, distributed systems, performance profiling, GPU utilization, PyTorch, ML platforms, Ads ML, ranking systems, recommender systems.
Practical notes
The base salary range for this position is $230,000 - $322,000 USD.
This role is eligible for equity in the form of restricted stock units.
Benefits include comprehensive healthcare, 401k with employer match, global benefits, family planning support, gender-affirming care, mental health benefits, flexible vacation, and generous paid parental leave.
Interviews may be recorded, transcribed, and summarized by AI, with an option to opt out. Personal information collected will be used for evaluation and will not be sold or shared for third-party marketing. Recordings will be deleted promptly after a hiring decision.
Reddit is an equal opportunity employer and is committed to providing reasonable accommodations for individuals with disabilities.