
Senior Engineer I/II, Drug Discovery Platform
Job description
About the role
You will architect and operate the data backbone that powers an AI-driven drug discovery factory closing the loop between computational design and wet-lab experiment in days instead of months. In this role, you own the molecule queue, the compound registration system, synthesis constraints, and the experimental result capture pipelines that feed back into predictive models. Your work will directly determine how fast Lila's AI learns from each cycle and therefore how fast new medicines are delivered. You will build the shared queue that brokers compounds between AI Scientists and the Make-Test platform, defining status state machines, batch grouping, priority semantics, and read/write contracts. You will also ensure data quality, lineage, schema evolution, and observability so that the closed-loop cycle time target of a few days from proposal to experimental truth is met reliably.
Key facts
What you'll do
Design and operate the shared molecule queue that brokers compounds between AI Scientists and the Make-Test platform, including status state machine logic, batch grouping rules, priority semantics, and robust read/write contracts.
Build the canonical molecule registry that handles SMILES and InChI normalization, stereochemistry handling, salt and parent resolution, duplicate detection, and stable internal identifiers such as LILA-NNN for downstream system reliability.
Construct pipelines that ingest data from ChemSpeed automated synthesis, QC instruments like LCMS and NMR, and bioassay readers, capturing identity, purity, dose-response curves, IC50 and EC50 values, ADMET measurements, and selectivity panels into a queryable store.
Maintain the live synthesis constraints and inventory state that the Batch Assembly AI reads each cycle, covering building-block availability, advanced precursor stock, ChemSpeed capacity, stock alerts, and chemistry-specific constraints such as reaction types, step budgets, and solvent compatibility.
Implement data quality controls, lineage tracking, schema evolution strategies, and observability dashboards to ensure the closed-loop cycle meets strict reliability targets and supports fast scientific iteration.
Collaborate with AI Scientists and experimental teams to translate wet-lab requirements into data models and API contracts that minimize latency and maximize feedback fidelity.
Partner with instrument and integration specialists to standardize data formats and workflows around ChemSpeed, LCMS, NMR, and bioassay systems for consistent, reusable pipelines.
Contribute to the design of a queryable store that supports complex scientific questions across compounds, batches, assays, and measurements with performance and scalability in mind.
Champion production-grade backend services in Python and SQL, ensuring robust error handling, monitoring, and maintainability across data platforms.
Adopt AI-assisted development tools such as Cursor and Claude Code, incorporating them effectively into day-to-day engineering work to accelerate delivery and code quality.
Apply strong database and schema design skills to evolving scientific data, working with structured, semi-structured, and unstructured data across relational and NoSQL stores.
Model chemical and biological data accurately, including structures, reactions, assays, dose-response experiments, batches, replicates, censored values, and failed runs to reflect real-world experimental noise.
Leverage cloud and containerized infrastructure on AWS with Kubernetes to deploy resilient, scalable services that meet the needs of a data-intensive discovery environment.
Engage with cheminformatics practices using tools like RDKit or equivalent for structure standardization, registration, and search to maintain a high-quality compound knowledge base.
Requirements
Hold a Bachelor's or Master's degree in Computer Science, Chemistry, Computational Biology, or a related field with a strong technical foundation.
Bring 5 or more years of experience building data platforms in production, ideally including exposure to scientific or lab-generated data and complex data integration challenges.
Demonstrate proven ability to design and ship data platform components from the ground up, including ingestion frameworks, registries, storage abstractions, and orchestration systems.
Write production-quality code in Python and SQL, with fluency in backend production APIs and services that are reliable, testable, and well-documented.
Show experience with database and schema design for evolving scientific data, optimizing queries across structured, semi-structured, and unstructured data sets.
Be comfortable modeling chemical and biological data, including structures, reactions, assays, dose-response curves, batches, and the messy realities of experimental measurements.
Have hands-on experience with AWS services and containerized deployment patterns using Kubernetes for scalable, observable systems.
Be proficient with AI-assisted development tools such as Cursor, Claude Code, or similar, and incorporate them effectively into engineering workflows to improve velocity and code quality.
Nice to have
Possess hands-on cheminformatics experience with RDKit, OpenEye, or equivalent for structure standardization, registration, and search.
Have direct experience building or operating a corporate compound registration system, such as CDD Vault, Dotmatics, Benchling Registry, or a similar platform, or building one from scratch.
Show familiarity with electronic lab notebooks, LIMS, and instrument data formats including mzML, AnIML, and vendor-specific formats.
Have background in workflow orchestration tools such as Flyte, Airflow, Dagster, or Temporal, especially for long-running scientific pipelines that integrate design, synthesis, and testing.
Practical notes
This is a full-time position based in Cambridge, Massachusetts, USA.
Employment is at-wisk.
Candidates must be authorized to work in the United States without sponsorship now or at the time of hire.
The role requires collaboration across AI, wet-lab engineering, and data teams with frequent interactions to align on priorities and data contracts.
Application review will proceed on a rolling basis until the position is filled.