
Staff Machine Learning Engineer, ML Efficiency
Job description
About the role
Reddit is a community of communities built on shared interests, passion, and trust, and it is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about, with 100,000+ active communities and approximately 130 million daily active unique visitors. The ML Efficiency team constructs the infrastructure, tooling, and optimization systems that enable machine learning engineers and researchers to train, evaluate, deploy, and operate models efficiently at scale. This role focuses on improving developer productivity, reducing infrastructure costs, increasing hardware utilization, and accelerating experimentation across the company's ML ecosystem. You will own the design and delivery of systems that make ML workflows faster, more reliable, and more cost-effective.
Key facts
What you'll do
Design and build systems that improve the efficiency of ML training and inference workloads across Reddit's global infrastructure.
Develop tooling that helps ML engineers debug, profile, optimize, and monitor model performance in production environments.
Improve GPU and general resource utilization through advanced scheduling, resource management, caching, and workload optimization strategies.
Partner with ML researchers and product teams to identify bottlenecks and drive measurable performance improvements.
Build benchmarking frameworks and performance dashboards that provide visibility into training and serving systems.
Optimize distributed training infrastructure, data pipelines, and model serving architectures for scalability and reliability.
Lead cross-functional initiatives that elevate the productivity of Reddit ML engineers and streamline platform adoption.
Drive technical strategy for ML platform scalability, reliability, and cost efficiency over the long term.
Implement observability and diagnostics capabilities that accelerate root cause analysis for training and inference issues.
Collaborate with infrastructure and platform teams to integrate ML efficiency solutions into Reddit's broader engineering ecosystem.
Champion best practices for performance engineering and contribute to open source projects where appropriate.
Continuously evaluate emerging hardware and software technologies to assess their impact on ML workload efficiency.
Translate complex performance data into actionable insights for both technical and non-technical stakeholders.
Mentor engineers on optimization techniques and foster a culture of efficiency across ML workflows.
Requirements
BS, MS, or PhD in Computer Science or a related field.
5+ years of software engineering experience with a strong track record of delivering production-grade systems.
Strong proficiency in Python for developing ML tooling, data pipelines, and analysis scripts.
Proficiency in at least one systems language such as Go, C++, Rust, or Java is preferred for high-performance components.
Experience building distributed systems at scale that handle large workloads and complex dependencies.
Experience with machine learning infrastructure, training systems, or model serving platforms is essential.
Deep understanding of performance engineering, systems optimization, and resource utilization metrics.
Strong debugging and profiling skills to diagnose issues in complex, distributed ML environments.
Ability to work effectively in a fast-paced, cross-functional environment with ambiguous problem spaces.
Strong written and verbal communication skills for collaborating with engineering teams and stakeholders.
Willingness to dive into legacy codebases and refactor critical paths to improve efficiency and maintainability.
Comfortable using version control, CI/CD systems, and cloud infrastructure tooling.
Commitment to writing high-quality, testable, and maintainable code that adheres to industry best practices.
Nice to have
Experience with large-scale recommendation, ranking, generative AI, or foundation model systems.
Experience with distributed training frameworks such as PyTorch Distributed, Ray, Tensorflow, Spark.
Familiarity with GPU architectures and performance analysis tools like NVIDIA Nsight and similar suites.
Experience optimizing cloud infrastructure costs across large ML workloads on public cloud platforms.
Contributions to internal platforms used by multiple ML teams across the organization.
Experience with building real-time ML inference applications and low-latency serving systems.
Knowledge of container orchestration platforms such as Kubernetes and workflow orchestration tools.
Understanding of data storage systems, including databases, data lakes, and streaming platforms used for ML workloads.
Practical notes
This role is remote from the United Kingdom.
Full-time engagement.
No specific compensation details are provided in the source.
No travel requirements, visa sponsorship, or application deadlines are specified in the source.