
Scientist II / Senior ML Scientist, Cofolding and Structure-Aware ML
Job description
About the role
You will train and evaluate cofolding models that jointly reason over proteins, ligands, binding context, and experimental readouts to improve AI-driven discovery decisions. This role owns the development of representation-learning approaches, including contrastive learning and self-supervised methods, to make DEL and related datasets more informative for binding and selectivity signals. You will collaborate directly with ML researchers, computational chemists, computational biophysicists, data engineers, and drug discovery teams to turn model outputs into physically and chemically meaningful hypotheses. The models you build must produce clear, interpretable signals that medicinal and computational chemists as well as biophysicists can interrogate and validate in downstream agent-driven workflows. You will design training objectives, evaluation frameworks, and data pipelines that connect molecular and protein context while rigorously guarding against dataset artifacts and leakage. This position focuses on making structured and unstructured biological data actionable for low-data learning scenarios and for guiding focused chemical space exploration. You will also work with research engineers to scale training, inference, and evaluation workflows and help expose model-derived capabilities as tools for scientists and AI agents.
Key facts
What you'll do
Train and evaluate cofolding models for protein-ligand and related molecular discovery applications using modern deep learning methods.
Employ contrastive learning, representation learning, self-supervised learning, and multimodal objectives to improve cofolding models trained on molecules, proteins, structures, and experimental readouts.
Develop modeling approaches that increase the usefulness of DEL data for learning binding, enrichment, selectivity, and structure-activity signals across chemical and biological spaces.
Build and evaluate models informed by Boltz, AlphaFold-style cofolding, equivariant GNNs, and related structure-aware machine learning frameworks.
Design training objectives, including contrastive, self-supervised, or multimodal objectives, that effectively connect ligands, proteins, structures, assays, simulations, and experimental data.
Construct rigorous evaluation frameworks that distinguish meaningful molecular learning from dataset artifacts, leakage, or spurious correlations.
Collaborate with data and platform teams to define datasets, labels, negatives, controls, and metadata required for robust model training.
Partner with computational chemistry and biophysics teams to translate model outputs into physically and chemically plausible hypotheses and experiments.
Work with low-data learning scientists to identify which DEL, assay, simulation, or structural data would most improve model performance in focused chemical spaces.
Work with research engineers to scale training, inference, and evaluation workflows for large scientific datasets and models.
Help expose trained models and model-derived capabilities as accessible tools for scientists and AI agents within discovery workflows.
Requirements
PhD or equivalent experience in machine learning, computational biology, computational chemistry, bioinformatics, computer science, or a related field with a strong publication record.
Hands-on experience training deep learning models for molecular, protein, structural biology, or scientific data applications using frameworks such as PyTorch or JAX.
Direct experience with contrastive learning, representation learning, self-supervised learning, or multimodal learning applied to scientific problems.
Familiarity with DEL or related selection, enrichment, screening, or molecular assay datasets and an understanding of how experimental design affects model learning.
Experience with protein-ligand modeling, cofolding, structure prediction, geometric deep learning, or structure-aware molecular machine learning methods.
Practical experience with PyTorch, JAX, or an equivalent ML framework, including distributed training patterns when relevant.
Ability to design careful experiments, ablations, and evaluations that are scientifically rigorous and reproducible for complex biological systems.
Strong understanding of data quality, leakage risks, negative construction, and benchmark design specific to molecular and structural datasets.
Demonstrated ability to collaborate across ML, data, computational science, and drug discovery functions in a fast-paced, interdisciplinary environment.
Nice to have
Hands-on experience with DEL data and prior work on protein-ligand or molecular optimization problems in drug discovery contexts.
Experience with Boltz, AlphaFold or AlphaFold-derived methods, equivariant GNNs, diffusion models, protein language models, or molecular encoders.
Experience training or extending cofolding, protein-ligand, protein-protein, structure prediction, diffusion, or geometric deep learning models at scale.
Experience with distributed model training and large-scale scientific data pipelines that integrate heterogeneous experimental sources.
Familiarity with active learning, closed-loop molecular design, or experiment-driven model iteration in discovery settings.
Experience integrating ML models with scientific tools, databases, and assay platforms to support downstream decision-making.
Practical notes
Full-time position based in Cambridge, MA, London, or San Francisco.
Relocation support may be available for qualified candidates based on location and role requirements.
This role requires authorization to work in the country of employment; sponsorship considerations will be evaluated per role and location.
Interviews will be conducted in English and may include technical assessments relevant to machine learning and scientific modeling.