Computational Biologist
Job description
Computational Biologist at Verge Genomics.
About the role
The will define and enable new product offerings by leveraging the company's proprietary human data and drug discovery engine. This role owns the development of cutting-edge computational methodologies that integrate multi-omic datasets to create predictive models for translational biology. You will lead high-impact projects that apply and adapt AI models to address challenges in disease biology, biomarker discovery, and target exploration. The position involves partnering with AI companies to co-develop next-generation foundation models specifically tailored for drug discovery. You will frame complex biological problems in computational terms and design solutions that are not only meaningful but also interpretable and experimentally testable. The role requires translating between biological domain knowledge and machine learning objectives to ensure scientific rigor and practical utility. You will also evaluate and create computational methodologies that support the generation of CONVERGE-powered insights for external partners.
Key facts
What you'll do
- Develop and evaluate cutting-edge computational methodologies integrating multi-omic datasets to develop predictive models for translational biology.
- Lead high-impact projects that apply and adapt AI models to translational challenges in disease biology, biomarker discovery, and target exploration.
- Lead partnerships with AI companies to co-develop next-generation foundation models for drug discovery.
- Frame biological problems in computational terms and design solutions that are biologically meaningful, interpretable, and experimentally testable.
- Design and implement evaluation methodologies for assessing AI model capabilities relevant to biological research and applications.
- Translate between biological domain knowledge and machine learning objectives to ensure alignment with project goals.
- Work with Verge's AI partners to deliver a best-in-class biology foundation model using Verge's proprietary datasets.
- Develop a novel approach that enables a powerful new product offering, such as patient stratification or biomarker discovery.
- Deliver at least two CONVERGE-powered insights projects to pharmaceutical and biotechnology companies.
- Build an internal agentic AI workflow that supports multi-modal biomedical reasoning and orchestration.
Requirements
Candidates must have a PhD in computational biology, AI/ML, applied statistics, biophysics, or an MS with professional experience in relevant fields. You must possess a minimum of five years of experience working in applied computational biology and the integration of multi-omic datasets, including RNA-seq, genotyping, and clinical data, with at least two years spent in a startup environment. Additionally, you need a minimum of two years of experience in relevant areas of translational science, demonstrating a deep understanding of target identification, biomarker discovery, and patient stratification. You must have a proven ability to implement, evaluate, and create computational methodologies that leverage machine learning, statistics, and AI for biological research and discovery. Fluency with state-of-the-art systems biology workflows, including off-the-shelf biological databases and computational biology tools, is required. You must have a track record of bridging biological domain knowledge with computational approaches to solve real scientific problems. A track record of individual innovation is essential, with published research or shipped work that has influenced pharmaceutical R&D decisions. You should have extensive experience running end-to-end RNA-Seq data analyses, covering quality control, read quantification, normalization, and biological interpretation. Excellent coding skills in Python are mandatory, along with experience in relevant machine learning and AI libraries such as PyTorch, HuggingFace, scikit-learn, pandas, and numpy. A demonstrable portfolio, such as GitHub repositories, research code, or shared notebooks, is highly preferred. Experience in building and evaluating machine learning models on biological data is required, particularly with transformer-based models like scGPT, Geneformer, ESM, or ProtBERT, and a deep understanding of feature selection and model interpretability. Professional experience with AI workflows, including natural language processing, retrieval-augmented generation, embeddings, vectorization of diverse data types, and working with large language models like GPT, is necessary. You must also have demonstrated experience with model evaluation and experimental design in a scientific context, including the setup of appropriate benchmarks and controls.
Practical notes
The engagement is full-time. The location is listed as San Francisco
Remote. No information regarding hours, travel, visa, application deadlines, or compensation is provided in the source material.