Product Manager, Data Ingestion & Quality
Job description
Product Manager, Data Ingestion & Quality at Protege
About the role
Protege is building the essential infrastructure for AI training data. This role focuses on ensuring the quality and usability of data as it enters our platform, transforming raw inputs into a trusted, catalog-ready state. You will own the product strategy for the data supply chain, making it repeatable and reliable across various data types and industries.
Key facts
What you'll do
Define the stages, validation checks, and quality standards for data flowing into our system, from initial receipt to catalog readiness.
Develop product requirements for systems that automatically extract and generate necessary metadata, such as transcripts, tags, and schema information.
Establish clear definitions of "catalog-ready" data and create the tools to enforce these standards, directly examining data to confirm compliance.
Collaborate with different industry teams to translate their specific data readiness needs into consistent, platform-wide standards.
Requirements
Possess 4-7 years of product management experience, with a primary focus on data pipelines, data quality systems, or data ingestion platforms.
Demonstrate strong technical understanding, including the ability to write SQL, interpret pipeline logs, identify schema inconsistencies, and grasp the nuances of data validation architectures.
Have experience managing products that ingest and standardize messy, inconsistent data from external partners.
Exhibit sound judgment in deciding whether to build new capabilities or partner with existing solutions, especially in rapidly evolving technical areas.
Possess the ability to communicate effectively with engineering teams and other product managers, holding both technical and product-focused discussions.
Nice to have
Experience with tools and frameworks for data quality, metadata management, or data cataloging, such as dbt, Great Expectations, or data contracts.
Familiarity with methods for de-identifying sensitive information, including Protected Health Information (PHI), Personally Identifiable Information (PII), or confidential business data.
A background in healthcare data operations, financial data infrastructure, or other fields where data accuracy is critical.
Exposure to machine learning training pipelines or AI data workflows.
Experience with data governance strategies.
Skills & tools
SQL
Data pipeline architecture
Data quality frameworks
Metadata generation
QA tooling
Cross-functional collaboration
Practical notes
This role is not intended for individuals whose primary experience is with customer-facing data products like analytics dashboards or BI tools. The focus is on the underlying data infrastructure.