AI Scientist, BioMedical AI
Job description
About the role
This role operates at the convergence of machine learning and single-cell biology to advance therapeutic discovery. The position centers on creating foundation models that extract insight from high-dimensional cellular data.
Data roles turn raw information into decisions. Analysts query databases and build dashboards. Data scientists build models that predict outcomes. Data engineers build the pipelines that move and store data. All three work closely with business teams and need a mix of statistics, coding, and communication. Nearly every modern company runs on data teams, from startups to banks. A strong portfolio of past analyses matters more than degrees in many hiring decisions.
Key facts
What you'll do
Foundation models are designed, trained, and refined for perturb-seq and other single-cell datasets to reveal biological mechanisms.
Large-scale single-cell and perturbation data undergo preprocessing, normalization, and integration through purpose-built pipelines.
Computational methods are aligned with experimental design and discovery goals in close collaboration with experimental biologists.
Multimodal integration is developed across transcriptomic, epigenomic, and proteomic single-cell readouts to unify biological evidence.
Methods and results are published in top scientific journals and presented at conferences to advance the broader research community.
Requirements
A PhD or equivalent in Computational Biology, Computer Science, Bioinformatics, or related discipline is required for this position.
Demonstrated expertise in single-cell data analysis, with a focus on perturb-seq, is a mandatory requirement for effective performance.
Hands-on experience with foundation models or large-scale self-supervised learning, such as transformers and variational autoencoders, is necessary.
Strong coding skills in Python and ML frameworks including PyTorch or TensorFlow, alongside single-cell analysis tools like Scanpy and Seurat, are required.
Experience in developing scalable ML methods for large biological datasets is essential given the volume and complexity of the data.
Nice to have
A background in causal inference, generative models, or representation learning for biological systems is viewed favorably.
Familiarity with multi-omics integration and cross-modal foundation models is considered an advantage for this role.
A track record of impactful publications in computational biology, machine learning, or related fields supports successful evaluation.
Practical notes
Xaira Therapeutics is an equal-opportunity employer committed to building a diverse and inclusive team.
This role is based in South San Francisco, California, and may involve work across San Francisco, Seattle, and London locations.
Typical interview steps
Data interviews commonly include a SQL or coding exercise, a statistics question, and a case study. Candidates may be asked to design a metric, interpret an experiment, or build a small model. Some companies give a take-home analysis. Expect questions about past projects and the business impact of your work. Interviewers often evaluate how you communicate uncertainty and business impact, not only the math. Bringing a clean write-up of a past analysis to the interview is well received.
Good to know
The role focuses on developing foundation models for biological and disease understanding.
Core tools include transformers, variational autoencoders, and other large-scale self-supervised learning methods.
The work targets historically hard-to-drug molecular targets and aims to improve drug development success.
Multi-omic data integration across transcriptomic, epigenomic, and proteomic modalities is central to the analysis.
The position contributes to publishing and presenting novel methods in scientific and academic communities.
Career growth
Data careers grow toward senior analyst, staff data scientist, or data engineering lead. Many professionals specialize in machine learning, analytics, or infrastructure. Cross-functional work with product and engineering teams becomes more important at senior levels. The field changes quickly, so continuous learning is part of the job. Professionals who can translate numbers into decisions tend to advance fastest.