Data Engineer, Platform
Job description
Data Engineer, Platform at Basis Research.
About the role
You will architect and maintain the core data infrastructure that powers Basis research and platform initiatives, translating ambiguous requirements into robust data products. You will own the design of scalable data pipelines that enforce strict quality, provenance, and governance from ingestion to consumption. This role requires you to act as a foundational builder, ensuring that data workflows are reliable, observable, and trustworthy for both internal researchers and external customers. You will collaborate deeply with cross-functional teams to prevent redundant effort and to establish shared datasets that serve multiple projects simultaneously. Your work will directly support the development of medium-scale models and the scaling of data operations as Basis serves a growing base of external users. You will be responsible for implementing data governance frameworks that uphold privacy, security, and compliance standards across all data flows. Ultimately, you will enable rigorous, reproducible research by constructing data foundations that stand the test of evolving models and real-world usage.
Key facts
What you'll do
Architect and implement data pipelines for model training, evaluation, and inference that handle scale and complexity with resilience.
Design and maintain feature stores and data platforms that serve multiple teams, reducing duplication and ensuring consistency across projects.
Develop comprehensive data quality frameworks and governance systems that enforce validation, anomaly detection, and lineage tracking.
Implement metadata management and documentation standards that make data assets transparent, discoverable, and trustworthy for all stakeholders.
Build and optimize data architectures for relational and NoSQL systems, balancing schema design, normalization, and performance tradeoffs.
Integrate and operate streaming and batch data processing workflows using distributed computing frameworks such as Spark and orchestration tools like Airflow or Dagster.
Manage data lifecycle processes including ingestion from object stores like S3, transformation, and delivery to data warehouses such as Snowflake or BigQuery.
Collaborate with research and platform teams to align data infrastructure with the needs of experiment lifecycle, model development, and reproducibility.
Implement solutions for data versioning and experiment tracking that support reproducibility and auditability across machine learning workflows.
Contribute to open-source data engineering tools such as Airflow, dbt, and Great Expectations to strengthen the broader ecosystem and internal practices.
Evaluate and adopt emerging data technologies including vector databases and embedding pipelines to support modern AI applications and reasoning systems.
Define and operationalize data governance, privacy, and security policies across pipelines, ensuring compliance and risk mitigation at scale.
Drive automation of data operations to reduce manual overhead, improve reliability, and enable teams to focus on high-value research problems.
Act as a technical leader in data infrastructure, mentoring others and elevating the standard of data quality and provenance across the organization.
Requirements
You must have demonstrated significant achievements in data engineering for ML/AI systems, with concrete examples of impact in production environments.
You must possess expert-level proficiency in SQL and strong capabilities in Python for data processing, ETL development, and automation.
You must have hands-on experience with distributed computing frameworks such as Spark or Dask and workflow orchestration tools like Airflow, Dagster, or Prefect.
You must have direct experience with cloud data platforms, including data warehouses such as Snowflake, BigQuery, or Redshift, as well as data lakes and object storage such as S3.
You must have deep knowledge of ML data requirements, including feature engineering, training/validation/test splits, data versioning, and experiment reproducibility.
You must understand data modeling principles for both relational and NoSQL systems, including schema design, normalization and denormalization tradeoffs, and performance optimization strategies.
You must value data provenance and documentation, ensuring that data pipelines are transparent, decisions are recorded, and others can trust and understand the data you deliver.
You must be capable of progressing with autonomy on complex data challenges, scoping data projects, making sound architectural decisions, and delivering complete solutions from ingestion through consumption.
You must be excited about enabling rigorous research through trustworthy data infrastructure that accelerates our ability to solve intractable problems at scale.
Nice to have
Experience with feature stores such as Tecton or Feast, or background in building and operating feature platforms.
Background in ML research or research engineering that provides deep understanding of data needs across the experiment lifecycle.
Experience with data lineage tools such as Apache Atlas, DataHub, or Monte Carlo for metadata management and observability.
Knowledge of vector databases and embedding pipelines for modern AI applications and reasoning systems.
Contributions to data engineering open-source projects including Airflow, dbt, or Great Expectations that demonstrate technical depth and community engagement.
Understanding of responsible AI principles and data governance practices that ensure ethical and compliant data usage.
Practical notes
The role is based in the New York Office and requires full-time on-site engagement.
Candidates must be eligible to work in the United States without sponsorship for this position.