[Remote] Data Scientist, AI Data Foundations
FlexBoardRemoteFull Time2d ago
PythonAzureLLMAIMLData EngineeringData ScienceComplianceSQLCommunityEngineeringPlatform
Job description
[Remote] Data Scientist, AI Data Foundations at FlexBoard.
About the role
FlexBoard is looking for a Data Scientist to join our AI Data Foundations team to build the infrastructure that powers our machine learning and AI applications. You will be responsible for creating high-quality data structures, managing feature stores, and designing graph databases to support our digital lending platform.
Key facts
What you'll do
- Architect and maintain vector stores for retrieval-augmented generation, including embedding pipelines, chunking, and indexing strategies.
- Manage feature store assets for offline and online inference, ensuring data freshness and point-in-time correctness.
- Develop graph data structures to map relationships between applicants, products, and lenders for analytical and AI use cases.
- Perform deep-dive data discovery on lending and behavioral datasets to identify trends, anomalies, and potential model drivers.
- Create curated datasets for AI consumption, ensuring proper documentation, governance, and data classification.
- Define evaluation frameworks to measure RAG retrieval performance, embedding quality, and graph completeness.
- Collaborate with ML engineers to ensure data structures effectively support model training and production workflows.
- Translate complex data findings into clear narratives for both technical and business stakeholders through dashboards and reports.
Requirements
- 4 to 7 years of experience in data science, ML engineering, or data roles, specifically focused on building data assets for model consumption.
- Hands-on experience designing vector stores for semantic search or RAG.
- Experience building or operating feature stores, including offline training and online serving patterns.
- Proficiency in graph database modeling and query writing.
- Strong technical skills in Python (pandas, NumPy, scikit-learn, PySpark) and SQL.
- Practical experience deploying embedding models and LLM tooling in production environments.
- Bachelor or Master degree in Computer Science, Statistics, Mathematics, Engineering, or a related quantitative field, or equivalent experience.
Nice to have
- Experience in FinTech, specifically with lending, deposits, credit, fraud, or KYC/AML data.
- Familiarity with AI/ML tools like MLflow, feature stores, and data catalogs.
- Knowledge of vector databases such as pgvector, Chroma, or FAISS.
- Experience with Azure data and AI services, including Azure AI Search and ADLS Gen2.
- Background in evaluating RAG systems using metrics like recall@k, faithfulness, and hallucination measurement.
- Exposure to graph algorithms such as centrality or community detection.
Skills & tools
- Python, SQL, PySpark, pandas, NumPy, scikit-learn
- Vector stores, Feature stores, Graph databases
- LLM tooling, Embedding models, RAG evaluation
- Azure data and AI services
Practical notes
FlexBoard has a history of providing H1B visa sponsorship, though this does not guarantee sponsorship for this specific position. The company is headquartered in Costa Mesa, California, and operates as a digital lending platform founded in 1998.