Senior Machine Learning Infrastructure Engineer, Embedding Platform
Job description
About the role
Reddit is a community of communities. It is built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet's largest sources of information. For more information, visit www.redditinc.com. The LS Embedding Machine Learning Platform team is at the forefront of building highly expressive, machine learning models that power Reddit's recommendation systems. We go beyond standard retrieval and ranking architectures, leveraging modern deep learning approaches and scalable model designs to enhance personalization across Reddit's ecosystem. Our work impacts content discovery, user engagement, and platform growth at a massive scale. As a Senior Machine Learning Infrastructure Engineer, you will work across both model development and ML platform to build large-scale learning systems that improve recommendation and personalization on Reddit. At the senior level, you will own major technical components end to end: designing models, implementing training and evaluation pipelines, and driving production deployment in close partnership with ML platform, product, and cross-functional ML teams.
Key facts
What you'll do
Design, train, and improve large-scale machine learning platforms for recommendation or personalization systems.
Own and deliver major ML systems components end to end, from problem framing through production rollout.
Build and optimize end-to-end ML pipelines spanning data preparation, feature generation, training, evaluation, and deployment.
Improve distributed training, model efficiency, and online inference performance.
Apply modern modeling approaches including sequence modeling and related foundation-model techniques to Reddit use cases.
Develop reliable serving and monitoring patterns for low-latency, high-throughput production ML systems.
Work with cross-functional partners across product, relevance, ads, and core ML teams to deliver measurable improvements in user experience and business impact.
Drive rigorous offline and online evaluation, including experimentation, model diagnostics, and feedback-loop improvement.
Contribute to engineering quality through strong code, design reviews, documentation, and operational excellence.
Champion the adoption of scalable model architectures and infrastructure best practices across the organization.
Collaborate closely with data scientists to translate research prototypes into robust, production-grade services.
Participate in on-call rotations to ensure high availability and rapid response to production incidents.
Mentor engineers and technical contributors to elevate the overall capability of the ML infrastructure organization.
Continuously explore emerging techniques and tools to future-proof Reddit's modeling and deployment stack.
Requirements
5+ years of experience in machine learning engineering, with a strong focus on large-scale ML infrastructure and recommendation or personalization systems.
Expertise in modern deep learning architectures, including sequence models and foundational models.
Experience building or scaling ML platform for large datasets and high-traffic production environments.
Demonstrated ability to independently scope and execute ambiguous technical work, while owning high-quality implementation details.
Solid understanding of distributed training and inference concepts, such as data parallelism, model parallelism, pipeline parallelism, or related optimization techniques.
Proficiency in Python and experience with modern ML frameworks such as PyTorch, TensorFlow, or similar.
Strong software engineering fundamentals, including system design, debugging, testing, and performance optimization.
Experience with A/B testing, model evaluation frameworks, and real-time feedback loops in large-scale production systems.
Excellent communication skills, with the ability to effectively present complex ideas to both technical and non-technical stakeholders.
Strong ownership mindset and comfort operating critical production services in a fast-paced environment.
Bachelor's degree in Computer Science, Engineering, or a related technical field or equivalent practical experience.
Ability to thrive in an environment of rapid change and high ambiguity while maintaining a bias for action.
Commitment to delivering high-quality results with attention to detail and reliability.
Nice to have
Experience with large-scale embedding systems and related optimization techniques.
Knowledge of recommendation system best practices and evaluation methodologies.
Familiarity with Reddit's tech stack and internal tools.
Experience contributing to open source ML projects.
Practical notes
This is a full-time position based in the United States. Remote work eligibility applies to locations within the United States. No specific compensation details are provided in this posting. The role may involve on-call responsibilities.