Forward Deployed Engineer, Infrastructure Specialist
Job description
About the role
You will own the full lifecycle of deploying Reducto's agentic document platform into customer environments, acting as the primary technical partner for enterprise infrastructure and platform teams. You will design and execute deployment strategies that respect strict hardware, networking, security, and reliability constraints without compromising performance or stability. This role requires you to be the hands-on operator and trusted advisor during some of the most critical infrastructure decisions our customers make. You will translate complex customer infrastructure requirements into robust, repeatable deployment patterns that de-risk adoption at scale. You will work directly with customer teams to debug live issues and ensure that Reducto operates seamlessly inside their existing environments. This position is not theoretical; it is a hands-on, high-visibility role that directly impacts customer success and product reliability. You will be responsible for codifying every deployment lesson learned into documentation and automation that empowers the entire organization.
Key facts
What you'll do
- Lead end-to-end deployment of Reducto into customer environments, including planning, configuration, testing, and rollout across diverse infrastructure constraints.
- Partner with enterprise IT, security, and platform teams to assess their current infrastructure, security posture, and data management practices to design tailored deployment strategies.
- Analyze and understand the hardware that powers our in-house ML models to ensure successful deployment onto customer infrastructure with optimal performance.
- Build monitoring, telemetry, and phone-home patterns that provide deep visibility into customer-controlled environments while strictly maintaining their security posture and compliance requirements.
- Debug and resolve customer-specific infrastructure issues end-to-end, navigating complex Kubernetes misconfigurations and cloud IAM edge cases to minimize downtime.
- Own incident response for customer deployments with world-class operational excellence, ensuring rapid resolution and clear communication during critical events.
- Codify successful deployment patterns, runbooks, and automation to transform one-off customer work into scalable infrastructure for future deployments.
- Collaborate with product and engineering teams to translate customer feedback and operational data into actionable improvements for the platform and deployment tooling.
- Maintain strict version control and configuration management discipline to ensure consistency, auditability, and rollback capability in every deployment.
- Evaluate and integrate emerging infrastructure technologies that can enhance the reliability, security, or performance of Reducto in customer environments.
- Serve as a technical mentor and escalation partner for internal support and operations teams, elevating the entire organization's ability to serve customers.
- Document every deployment scenario, constraint, and outcome to build a durable knowledge base that accelerates future engagements.
- Conduct regular reviews of deployment logs, monitoring data, and customer feedback to identify patterns and drive systemic improvements.
- Champion best practices for security, reliability, and performance across all customer infrastructure touchpoints.
Requirements
- Are your own worst critic - have an extremely high bar for quality and always aim for robust solutions rather than quick fixes.
- Have 5+ years of hands-on experience operating production infrastructure, with meaningful time spent deploying software into environments you don't fully control.
- Are comfortable with Python or similar languages, and exceptional at working across cloud platforms, container orchestration (e.g., Kubernetes), networking, and storage technologies.
- Communicate clearly with external engineers. You can sit in a call with a customer's platform team, diagnose a problem live, and come out with a path forward that both sides trust.
- Are energized by being assigned to a customer problem and owning it end-to-end from initial assessment through resolution and documentation.
- Build your own tools on the fly to diagnose, experiment, and address reliability problems - whether it's an internal dashboard or an automated remediation workflow.
- Bring a quantitative, hands-on approach to system operations, automation, and continuous improvement.
- Thrive in ambiguous, fast-paced environments common to early-stage companies where priorities shift quickly and adaptability is essential.
Nice to have
- Have deployed software into regulated environments (financial services, healthcare, legal, insurance) with complex infrastructure.
- Have worked with Replicated/KOTS, Helm, or similar enterprise distribution tooling.
- Have experience with GPU infrastructure, model serving, or AI/ML workload deployment patterns on customer-controlled hardware.
- Have prior experience at an early-stage, high-growth company - as a founder, founding engineer, or early deployment hire.
Practical notes
This is an in person role at our office in SF. We're an early stage company which means that the role requires working hard and moving quickly. Please only apply if that excites you.