Principal Scientist, Machine Learning
Job description
About the role
Flagship Pioneering is seeking a Principal Scientist in Machine Learning to join a team that applies advanced computational methods to transform how biological systems are understood and how therapeutic candidates are identified. This role sits at the intersection of machine learning research and biological discovery, where the work directly shapes the direction of new ventures and platforms within the Flagship ecosystem. The Principal Scientist will lead model development efforts, collaborate with cross-functional teams of biologists, chemists, and engineers, and contribute to the publication of findings that advance the field. The position requires someone who can independently drive research agendas while mentoring junior scientists and communicating complex technical concepts to both technical and non-technical stakeholders. This is a hands-on leadership role that demands deep expertise in machine learning architectures and a genuine curiosity about biological systems.
Key facts
What you'll do
Design and implement deep learning models for genomics and multi-omics data analysis to uncover patterns relevant to drug discovery and therapeutic development.
Develop novel machine learning architectures tailored to biological sequence data, including protein sequences, nucleic acid structures, and gene expression profiles.
Lead the training and evaluation of large-scale models using high-performance computing infrastructure and distributed training frameworks.
Collaborate with experimental biologists and computational scientists to define research problems, translate biological questions into machine learning tasks, and validate model predictions against wet-lab results.
Mentor and guide a team of machine learning engineers and scientists, providing technical direction and code review to ensure research quality and reproducibility.
Publish research findings in peer-reviewed journals and present results at scientific conferences to establish Flagship Pioneering as a leader in ML-driven biology.
Build and maintain reusable software libraries and pipelines that streamline the end-to-end workflow from raw biological data to trained models and interpretable results.
Partner with venture creation teams to evaluate emerging opportunities where machine learning can address unmet needs in healthcare and life sciences.
Conduct literature reviews and stay current with advances in machine learning, computational biology, and related disciplines to identify new methodologies that can be applied to Flagship projects.
Communicate research progress and technical insights to leadership and external collaborators through written reports, presentations, and strategic discussions.
Requirements
PhD in Computer Science, Machine Learning, Computational Biology, Bioinformatics, or a closely related field with a strong focus on machine learning.
8 or more years of experience applying machine learning methods to complex biological or biomedical datasets in an academic or industry research setting.
Demonstrated expertise in deep learning frameworks such as PyTorch or TensorFlow, with hands-on experience training and deploying models at scale.
Strong background in genomics, transcriptomics, proteomics, or a related omics discipline, with the ability to work fluently across biological and computational domains.
Proven track record of first-author publications in top-tier venues such as NeurIPS, ICML, ICLR, Nature Methods, Genome Research, or comparable journals.
Experience with version control, containerization, and cloud-based computing environments, including AWS or GCP, for managing large-scale data processing and model training workflows.
Nice to have
Experience with transformer-based architectures, graph neural networks, or attention mechanisms applied to biological sequence or structure prediction tasks.
Familiarity with single-cell sequencing data and methods for analyzing cellular heterogeneity using machine learning approaches.
Prior experience in a startup or venture creation environment, where rapid iteration and cross-disciplinary collaboration are essential to success.
Knowledge of drug discovery pipelines, including target identification, hit discovery, and lead optimization, and how machine learning can accelerate each stage.
Skills & tools
Python programming for data analysis, model development, and production-grade software engineering in scientific computing contexts.
Deep learning frameworks including PyTorch and TensorFlow for building, training, and evaluating neural network models on biological datasets.
Genomics and bioinformatics toolkits such as Biopython, Scanpy, Seurat, or equivalent libraries for processing and analyzing high-dimensional biological data.
Cloud computing platforms including AWS, Google Cloud Platform, or Azure for scalable data storage, distributed training, and model deployment.
Version control systems such as Git and collaborative development workflows using GitHub or GitLab for managing research code and reproducibility.
Data visualization and analysis libraries including NumPy, pandas, scikit-learn, and matplotlib for exploratory data analysis and model interpretation.
Practical notes
This role is based in our Cambridge, MA office and requires on-site presence during standard business hours, with flexibility for hybrid arrangements as determined by team needs.
The Principal Scientist will work within a multidisciplinary team that includes biologists, chemists, data engineers, and software developers, requiring strong communication skills and the ability to translate between disciplines.
Flagship Pioneering is an equal opportunity employer and welcomes applicants from all backgrounds. The company values diversity of thought and experience as essential to scientific innovation.
Candidates should be prepared to discuss specific past projects, including the datasets used, models developed, and biological insights gained, during the interview process.