
ML Scientist I / II, Foundation Models for Life Sciences
Job description
About the role
This position focuses on exploratory research and the translation of biological questions into model behavior. The role owns end-to-end scientific problems in generative biology, guiding model design and closing the experimental loop. You will collaborate with wet-lab teams to contribute to Lila methods and the organization's presence in the scientific community. The position is an IC role for individuals building deep expertise in generative AI applied to biology. You will drive the definition of research objectives that connect biological intuition with model-driven hypotheses. Your work will establish the criteria for evaluating biological relevance and scientific rigor in model outputs. You will iterate on model behavior based on experimental feedback to ensure alignment with real-world biological constraints. This role bridges quantitative modeling and experimental validation to advance the state of the art in generative life sciences.
What you will do
You will contribute to research on foundation models for life science applications, including biological sequence design, structure prediction, and multimodal scientific reasoning. You will design, train, and evaluate generative models on biological and chemical data while incorporating domain-specific constraints and priors into the modeling process. You will be part of the end-to-end ML process within Lila's "Lab-in-the-Loop" lifecycle, supporting data generation strategy, building pipeline models, and helping design feedback loops where experimental results improve model performance. You will translate ambiguous biological questions into well-defined ML problems and interpret model outputs in close collaboration with wet-lab scientists and computational biologists. The role involves establishing and maintaining research quality and methodology standards across the foundation models program. You will create and refine feedback mechanisms where experimental outcomes directly inform model architectures and loss function design. Translating domain constraints into model inductive biases for protein and nucleic acid problems is a core duty of this position. Publishing and presenting methods that strengthen Lila's presence in scientific communities is an expected component of this role. You will work with high-dimensional biological data to identify emergent patterns that inform scalable model training strategies and data efficiency. You will evaluate model uncertainty and robustness in the context of biological risk and experimental feasibility to ensure safe deployment.
Requirements
You must hold a PhD or equivalent research experience in a quantitative field within two years of joining this role. A strong foundation in generative model architectures and training methodologies is required, with hands-on experience in model development, tuning, and rigorous evaluation. You must be able to own research projects from initial problem framing through experimentation, analysis, and iteration independently with minimal supervision. Familiarity with at least one life science domain such as genomics, protein engineering, or molecular design is necessary for domain-specific reasoning. Experience collaborating with experimental teams or working directly with biological datasets is a strict requirement for successful integration. Proficiency in PyTorch, JAX, or TensorFlow and associated GPU-based training workflows is mandatory for daily operations. You must demonstrate rigorous experimental practices, including systematic ablation studies and controlled comparisons to ensure scientific validity. Comfort with version control, reproducible pipelines, and experiment tracking is essential for maintaining research integrity and facilitating collaboration. You should be able to communicate complex model behavior and limitations to interdisciplinary audiences clearly and constructively. The ability to rapidly prototype ideas and translate theoretical concepts into testable hypotheses is critical for this position.
Nice to have
Experience in computational protein design or molecular structure prediction methods is valued and aligns with current team needs. Exposure to active learning or closed-loop experimental optimization pipelines is beneficial for advancing autonomous research capabilities. Contributions to open-source scientific ML tools or curated benchmark datasets are recognized as signals of community engagement. Familiarity with distributed training infrastructure at scale is an advantage for handling large biological datasets efficiently. High-impact publications in AI for Science venues or related conferences are considered bonus points for research impact.