Machine Learning Infrastructure Tech Lead
Job description
About the role
Reducto is an agentic document platform providing specialized tools for document processing by combining in-house and frontier models. We are seeking a hands-on technical leader to manage our training and inference systems as we scale our document workflows for enterprise clients. The Machine Learning Infrastructure Tech Lead will own the end to end lifecycle of our machine learning infrastructure, ensuring that our systems are robust, scalable, and aligned with enterprise demands. This role involves driving technical strategy while maintaining a sharp focus on performance, reliability, and efficiency across document processing workflows. You will be responsible for translating high level product goals into concrete infrastructure capabilities that enable rapid experimentation and dependable deployment. The ideal candidate will balance deep technical execution with leadership, guiding the team to deliver infrastructure that unblocks machine learning innovation. You will partner closely with research and product teams to ensure our infrastructure evolves in step with model advancements and business needs.
Key facts
What you'll do
- Define the technical roadmap and strategy for our machine learning infrastructure, setting clear priorities for capacity and capability growth.
- Develop and support stacks for model training and inference that balance rapid iteration with production performance, ensuring stability at every stage.
- Improve serving efficiency by optimizing kernels, runtimes, batching, scheduling, and distributed inference to extract maximum throughput from existing hardware.
- Architect systems for reliable multi-node and multi-GPU training and serving, designing for fault tolerance, scalability, and operational simplicity.
- Enhance GPU utilization, throughput, latency, cost efficiency, and observability, creating clear metrics and dashboards to guide investment decisions.
- Create benchmarks to pinpoint performance bottlenecks and prioritize infrastructure investments, aligning engineering effort with the highest impact opportunities.
- Integrate relevant state-of-the-art advances in training and inference into our systems, evaluating new techniques for practical benefit and scalability.
- Create abstractions and tooling to accelerate the transition of models from research to production, reducing friction for data scientists and engineers.
- Collaborate with platform and ML teams on capacity planning and architectural design, ensuring infrastructure decisions support long term product goals.
- Mentor engineers and lead design reviews to maintain high technical standards, fostering a culture of quality, clarity, and continuous improvement.
Requirements
- 5+ years of experience in production infrastructure, with a focus on ML systems, demonstrating a history of designing and operating reliable services.
- Proven track record of leading ambiguous technical projects through to deployment, showing comfort with uncertainty and ownership of outcomes.
- Proficiency in Python and systems engineering, with the ability to write clean, maintainable, and performant code.
- Deep understanding of performance characteristics for GPU training and inference workloads, including memory bandwidth, compute patterns, and latency tradeoffs.
- Experience with Kubernetes and distributed training or serving frameworks, including orchestration, scheduling, and resource management.
- Ability to bridge the gap between low-level model performance and platform architecture, translating research needs into production constraints.
- Commitment to operational reliability and high-quality engineering, with a mindset focused on monitoring, alerting, and incident response.
- Strong communication skills to collaborate effectively with cross functional stakeholders and translate technical details into actionable plans.
Nice to have
- Experience optimizing or writing custom model-serving kernels, Triton, or CUDA, demonstrating a willingness to dive into low level performance work.
- Contributions to open-source projects like vLLM, SGLang, PyTorch, TensorRT-LLM, or Ray, showing engagement with the broader community and proven technical depth.
- Background operating distributed training or inference across hundreds or thousands of GPUs, with evidence of scaling complex workloads.
- Experience building scheduling, capacity management, or observability systems for GPU workloads, including custom tooling and dashboards.
- Prior experience in high-growth or early-stage startup environments, with adaptability to fast changing priorities and resource constraints.
Practical notes
This is a fully in-person role based in San Francisco. Benefits include unlimited PTO, daily catered lunches, commuter reimbursement, comprehensive medical/dental/vision insurance, a $150 monthly health and wellness budget, and flexible parental leave. Reducto is an equal opportunity employer.