Data Product Operations Lead
Job description
About the Role
David AI operates at the intersection of audio and artificial intelligence, establishing the first research company dedicated exclusively to audio data. We have built the Data Factory, a proprietary infrastructure designed to transform raw audio into high-quality training datasets with the rigor typically reserved for model development. Our mission is to integrate AI into the tangible world, and we believe speech is the most accessible and human-like gateway to this integration. As the demand for novel audio use cases grows, the bottleneck shifts to high-quality data; our team exists to solve that problem. This Data Product Operations Lead role is central to our thesis, bridging the gap between research ambition and operational reality.
This position is based in San Francisco on a full-time basis. You will join a team described as sharp, humble, ambitious, and tight-knit, working alongside former leaders from the Scale AI ecosystem. In less than a year, David AI has secured most FAANG companies and AI labs as customers, backed by a $50M Series B from prominent investors including Meritech, NVIDIA, Jack Altman, Amplify Partners, and First Round Capital.
What You'll Do
You will own the complete lifecycle of data products that power the core Data Factory at David AI. This involves translating ambiguous research ambitions from leading AI labs into concrete, scalable data pipelines that generate high-quality audio. You will act as the operational backbone of our R&D process, ensuring data workflows progress from 0→1 prototypes to 1→N production systems without losing fidelity.
Specifically, you will orchestrate end-to-end data pipelines, designing and implementing workflows that process high-volume audio inputs while maintaining strict quality, reliability, and efficiency thresholds. You will partner directly with researchers to deconstruct emerging model capabilities and convert them into actionable data collection and validation strategies. Your role requires leading cross-functional initiatives that align Operations, Product, and Engineering teams around unified data factory objectives.
You will establish monitoring frameworks to track pipeline health, identifying and resolving issues related to sourcing, quality, or process inefficiencies. Using metrics such as throughput, quality, and cost, you will drive prioritization and continuous improvement. This includes taking full ownership of customer outcomes, diagnosing bottlenecks, and architecting solutions to ensure on-time dataset delivery. You will translate ambiguous product requirements into concrete technical specifications and implement robust testing and validation protocols to guarantee audio data integrity throughout the production lifecycle.
Additional responsibilities include championing best practices across the data lifecycle, from raw ingestion to final dataset delivery. You will drive process optimization by identifying and automating repetitive tasks to increase team efficiency. This role requires synthesizing qualitative researcher feedback into quantitative pipeline improvements and mentoring junior operations staff to elevate the entire Data Operations organization. You will act as the primary point of contact for internal stakeholders regarding data pipeline performance and roadmap alignment.
Requirements
The ideal candidate for this Data Product Operations Lead role brings 1-6 years of experience in high-intensity environments. This background may include founder experience, strategy consulting, venture-backed operational roles, investment banking, or similar fields. A background in computer science or a closely related technical field is also listed as a valid path.
You must demonstrate systems thinking, with the ability to design for scale, durability, and long-term maintainability. Strong product intuition is essential for effective collaboration with engineers and researchers to solve problems quickly. The role requires a high-execution operator who is fast, detail-oriented, and uncompromising on the quality of data outputs.
You must thrive in collaborative settings, maintaining a low ego and a willingness to perform hands-on work across the operational spectrum. Resilience is required to navigate ambiguity and build clarity from vague or evolving project scopes. A commitment to maintaining the highest standards of data integrity, security, and ethical handling is non-negotiable.
Bonus Points
The listing specifies that a track record of extreme ownership is a significant advantage, defined as focusing on outcomes rather than merely completing tasks. Prior experience in data, machine learning, or large-scale production operations environments is also noted as beneficial.