Don't See a Perfect Role? Apply Anyway!
Job description
Reducto Document Intelligence Engineer at Reducto.
About the role
This role enables AI teams to operate at enterprise scale using document intelligence. The position supports critical document workflows for customers ranging from AI labs to large enterprises.
Data roles turn raw information into decisions. Analysts query databases and build dashboards. Data scientists build models that predict outcomes. Data engineers build the pipelines that move and store data. All three work closely with business teams and need a mix of statistics, coding, and communication. Nearly every modern company runs on data teams, from startups to banks. A strong portfolio of past analyses matters more than degrees in many hiring decisions.
Key facts
What you'll do
Analyze complex enterprise documents and define processing workflows that turn unstructured files into reliable data for AI systems. This work targets AI teams and reduces manual effort in document processing.
Build extraction and parsing capabilities that handle tables, charts, and nonstandard layouts using vision models and LLMs combined with heuristic rules. These capabilities convert documents into reliable data for downstream applications.
Design document chunking and metadata strategies that power accurate retrieval and context-aware downstream applications. The strategies use document layout data to improve context preservation.
Translate ambiguous business requirements into robust document processing pipelines that reduce manual effort and errors. The pipelines address complex documents and aim to lower operational overhead.
Partner with product and engineering teams to integrate extraction and search workflows into production services. Integration ensures that document workflows align with product goals.
Support high-volume use cases for customers such as insurance, finance, and enterprise operations by ensuring reliability and scalability. Reliability and scalability serve customer demands in insurance, finance, and enterprise operations.
Create clear documentation of processing steps and data transformations so that solutions are maintainable and auditable. Documentation helps teams understand and validate document processing flows.
Requirements
Demonstrated experience designing or implementing document processing or information extraction systems. Experience ensures that systems can handle real-world document complexity.
Strong proficiency in at least one general-purpose programming language for building production-grade data pipelines. Proficiency supports reliable extraction and parsing logic.
Familiarity with LLMs, vision models, or structured extraction techniques for turning documents into reliable data. Familiarity enables effective system design.
Experience collaborating with cross-functional teams including product, engineering, and customer-facing stakeholders. Collaboration aligns workflows with product and business needs.
Ability to work effectively here involves handling evolving priorities and adapting to customer needs and product direction. Adaptability supports success in ambiguous problem spaces.
Willingness to engage with complex, real-world documents that require iterative experimentation and robust validation. Validation ensures reliability across diverse document types.
Practical notes
This role is based in the San Francisco Office and follows an onsite work model. The position is full-time and eligible for the benefits outlined in company policies. Typical interview steps
Data interviews commonly include a SQL or coding exercise, a statistics question, and a case study. Candidates may be asked to design a metric, interpret an experiment, or build a small model. Some companies give a take-home analysis. Expect questions about past projects and the business impact of your work. Interviewers often evaluate how you communicate uncertainty and business impact, not only the math. Bringing a clean write-up of a past analysis to the interview is well received.
Good to know
Document intelligence combines vision, language models, and rule-based heuristics to convert unstructured files into structured data. Modern retrieval systems rely on careful chunking and metadata design to preserve context. Production pipelines must balance accuracy, latency, and maintainability across diverse document types. Cross-functional collaboration is common when integrating extraction workflows into larger products. Rapid iteration and validation are necessary when handling complex, real-world document formats.