Machine Learning Infra Engineer
Job description
About the role
You will own the design and operation of the production-grade systems that power Reducto's document AI at enterprise scale. This role centers on building the inference and training frameworks that turn complex document workflows into reliable, high-throughput services. You will partner closely with ML researchers and platform engineers to turn novel methods into performant, observable systems. You will set the bar for how large language models and vision models are served and trained across GPU clusters. This is a hands-on role where your work will directly determine the speed and reliability of product delivery. You will take full ownership of reliability, performance, and cost across the entire ML infrastructure stack. Your contributions will enable Reducto to serve hundreds of enterprise customers with strict performance and uptime requirements.
Key facts
What you'll do
- Design and maintain the end-to-end training and inference stack, emphasizing fast iteration and flexibility for new methods alongside high-performance serving.
- Create and evolve benchmarks for training and inference pipelines to pinpoint bottlenecks and drive targeted optimizations.
- Evaluate and integrate state-of-the-art advances in training and inference to keep Reducto at the forefront of document AI performance.
- Architect systems for scaling model training across multi-node, multi-GPU environments with strict reliability and rich observability.
- Scale distributed training and inference workloads across large GPU clusters while improving utilization, reliability, and cost efficiency.
- Build intuitive tooling, abstractions, and observability that allow ML engineers to move rapidly from experiment to production.
- Implement robust data loading, preprocessing, and caching strategies to maximize GPU utilization and minimize idle time.
- Collaborate with platform and ML teams to standardize deployment patterns, monitoring, and alerting across services.
- Diagnose and resolve performance regressions in production inference paths using deep system-level insight.
- Define and own reliability targets, failure modes, and rollback strategies for critical inference services.
Requirements
- Hold a high bar for quality, precision, and consistency in all engineering deliverables.
- Enjoy solving complex, ambiguous problems by breaking them down and building from first principles.
- Bring strong Python expertise and a solid background in systems engineering and software design.
- Demonstrate comfort with Kubernetes, containerization, and distributed training frameworks such as PyTorch or JAX.
- Have hands-on experience deploying and operating services in production cloud environments.
- Thrive in fast-paced, high-growth startup environments where priorities evolve quickly.
- Collaborate effectively with both technical and non-technical stakeholders across teams and roles.
- Take full ownership from initial strategy and requirements through design, implementation, and on-call responsibilities.
- Have a minimum of 3 years of professional experience building and operating software systems.
Nice to have
- Have experience at an early-stage or high-growth startup where systems must scale quickly.
- Have contributed meaningfully to open source projects in training or inference stacks.
- Have hands-on experience setting up distributed inference across hundreds to thousands of GPUs.
- Show deep passion for combining technical excellence with measurable business impact.
Practical notes
This is an in-person role at our office in San Francisco. We are an early-stage company, which means the role requires working hard and moving quickly. Please only apply if that environment excites you and aligns with your career goals.
More about Reducto
Nearly 80% of enterprise data is in unstructured formats like PDFs.
PDFs are the status quo for enterprise knowledge in nearly every industry. Insurance claims, financial statements, invoices, and health records are all stored in a structure that's simply impractical for use in digital workflows. This creates a critical bottleneck that leads to dozens of wasted hours every week https://www.reducto.ai/blog/the-real-cost-of-manual-document-processing.
Traditional approaches fail at reliably extracting information in complex PDFs.
OCR and even more sophisticated ML approaches work for simple text documents but are unreliable for anything more complex. Text from different columns are jumbled together, figures are ignored, and tables are a nightmare to get right. Overcoming this usually requires a large engineering effort dedicated to building specialized pipelines for every document type you work with.
Reducto breaks document layouts into subsections and then contextually parses each depending on the type of content. This is made possible by a combination of vision models, LLMs, and a suite of heuristics we built over time. Put simply, we can help you accurately extract text and tables even with nonstandard layouts, automatically convert graphs to tabular data and summarize images in documents, extract important fields from complex forms with simple natural language instructions, build powerful retrieval pipelines using Reducto's document metadata, and intelligently chunk information using the document's layout data.
Benefits at Reducto
At Reducto, we're invested in the well-being and growth of our team. Here's what we currently offer:
- Unlimited PTO: We believe great work requires recharging.
- Lunch: Receive a free lunch to eat with your teammates daily at the office
- Reimbursed Transportation: Provide us with your receipts and we'll take care of the costs
- Insurance: Generous health, dental, and vision coverage for you and your dependents
- Equity: Competitive stock options to align your success with the company's growth
- Mental Health Support: Access to counseling and therapy resources
- Learning Budget: Funds for courses, conferences, and books to support your growth
- Flexible Schedule: The freedom to manage your time around core collaboration hours
- Remote Support: Stipends for home office equipment as needed
- Community: Regular team offsites and social events to build strong relationships
This is an in-person role at our office in San Francisco. We are an early-stage company which means that the role requires working hard and moving quickly. Please only apply if that excites you.