Head of Data
Job description
About the role
The individual in this position will own the end to end lifecycle of data assets within a high growth AI company focused on materials discovery. They will define and drive the data strategy that connects advanced research in computational materials science with scalable production grade data pipelines. This role requires close partnership with scientific founders and domain experts to translate complex experimental and simulation outputs into structured, high quality training data. The Head of Data will establish robust data governance, ensuring that data quality, provenance, and security standards are met across the entire organization. They will lead the design and deployment of data infrastructure that supports both exploratory research and long term product development. The successful candidate will build and mentor a growing data team, fostering a culture of rigor, curiosity, and cross functional collaboration. They will play a key role in aligning data initiatives with business milestones and scientific validation efforts. This position offers the opportunity to shape the foundational data layer of a company aiming to transform materials R and D through artificial intelligence.
Key facts
Location: Singapore
Engagement: Full-time, Hybrid (3 days in the office).
Some travel may be necessary for collaborative efforts. Benefits include 21 days of vacation (Singapore), parental leave (26 weeks for primary caregivers and 12 weeks for secondary caregivers), and a budget for professional development. CuspAI is committed to being an equal opportunity employer and encourages applications from individuals of diverse backgrounds. Accommodations for the interview process can be arranged upon request.
What you'll do
Architect and maintain the central data lake and warehouse, integrating heterogeneous sources such as simulation outputs, experimental measurements, and literature derived datasets.
Lead the design and implementation of data pipelines that ensure timely, reliable, and reproducible movement, transformation, and storage of structured and unstructured data.
Establish data quality frameworks and validation checks to monitor integrity, consistency, and freshness across all datasets used for modeling.
Define and enforce metadata standards, data dictionaries, and versioning practices to enable clear provenance and reproducibility for scientific and analytical use cases.
Partner with machine learning and scientific teams to identify data requirements for new modeling projects and to support the generation of high impact training sets.
Implement data access controls and privacy preserving mechanisms, balancing open scientific collaboration with the protection of proprietary and sensitive information.
Drive the adoption of data visualization and exploration tools, enabling researchers and engineers to derive insights quickly and make informed decisions.
Oversee the curation of specialized datasets for materials properties, reaction conditions, and characterization results, ensuring they are well documented and easily discoverable.
Champion the use of automated testing and monitoring for data pipelines, detecting anomalies, drift, and failures before they impact downstream applications.
Guide the selection and evaluation of data storage technologies, workflow orchestration tools, and infrastructure that scale with the growth of the data estate.
Develop and deliver training for cross functional teams on data best practices, tools, and standards to improve overall data literacy.
Lead the roadmap for data product development, translating scientific needs into concrete data deliverables that accelerate research and commercialization.
Coordinate with external collaborators and academic partners on data sharing agreements, ensuring compliance with legal, ethical, and institutional requirements.
Continuously evaluate emerging data technologies and methodologies, proposing pilots or upgrades that enhance capability and efficiency.
Requirements
You hold a Master's or PhD degree in a quantitative field such as computer science, statistics, mathematics, physics, or computational chemistry.
You have demonstrated experience building and operating data platforms, preferably in a cloud native environment using tools like Snowflake, BigQuery, or similar.
You possess a strong background in software engineering, with proficiency in Python, SQL, and at least one modern data stack framework.
You have managed data pipelines and ETL workflows at scale, ensuring reliability, performance, and maintainability in production settings.
You understand the principles of data governance, including metadata management, lineage tracking, and access control policies.
You are comfortable working with large, complex, and often noisy scientific datasets, particularly those originating from simulations or experiments.
You have a track record of collaborating with interdisciplinary teams, translating ambiguous problems into concrete analytical deliverables.
You hold the right to work in Singapore or one of the alternative locations listed, with eligibility to obtain necessary work visas if required.
Nice to have
Experience in the materials science, chemistry, or drug discovery domains is highly valued.
Familiarity with specialized data formats and tools used in computational chemistry or high throughput experimentation.
Knowledge of containerization, orchestration, and infrastructure as code practices for data platforms.
Practical notes
This is a hybrid role requiring 3 days in the office in the preferred location.
Some travel may be required for collaborative projects.
Candidates must meet the eligibility requirements related to work rights and visa sponsorship as noted in the requirements section.