Member of Technical Staff
Job description
About the role
We are looking for a member of technical staff focused on data collection and research to own how LatchBio sources, curates, and packages data. In this role you will build relationships with biotech companies, contract research organizations, and data clearinghouses to source and generate unique, high-value datasets. The datasets you create will be used to train agents on real-world research and development tasks across biology. You will take ownership of use case scoping and the data curation strategy under three verticals: multi-omics, therapeutics, and biosecurity. This is a hands-on position that sits at the intersection of bioinformatics, partnerships, and product thinking. You will spend your time both talking to data providers and working directly with raw data to turn it into something useful for the rest of the organization.
Key facts
- Employer: LatchBio
-
Location: USA
-
Job type: Full time
- Focus areas: Multi-omics, therapeutics, and biosecurity data
- Core responsibility: Data sourcing, curation, and benchmarking support
- Main function: Member of technical staff for data collection and research
What you'll do
- Build partnerships with data vendors, institutions, and laboratories to secure access to proprietary data
- Own the end to end pipeline for sourcing, cleaning, and packaging data used for biology benchmarks
- Scope new use cases and decide which datasets matter most across the three verticals
- Source clinical and molecular datasets directly from institutions and research groups
- Assemble and package public data such as patents, publications, molecular atlases, and clinical trials into realistic biological analysis tasks
- Use bioinformatics tools to process raw datasets into formats ready for benchmarking
- Work ahead of the bench, staging data before it is needed rather than after
- Collaborate with research engineers and product teams to define what a usable dataset looks like
- Maintain documentation for data provenance, licensing, and quality controls
- Track data inventory and gaps, and propose new sources to close those gaps
- Represent LatchBio in conversations with external data providers and academic partners
Requirements
- Strong background in computational biology, bioinformatics, or a related field
- Experience working with omics data such as genomics, transcriptomics, or proteomics
- Familiarity with clinical trial data and molecular datasets
- Ability to write and run scripts for data processing and cleanup
- Understanding of data licensing and institutional agreements
- Comfort working independently and owning a domain end to end
- Strong communication skills for working with external partners
- Demonstrated interest in training data quality and AI evaluation
Nice to have
- Experience at a startup or on a data team inside a life sciences company
- Prior work with patent databases or scientific literature mining
- Experience with biosecurity datasets or related policy work
- Familiarity with public molecular atlases and reference resources
Skills & tools
- Bioinformatics tooling and common bioinformatics libraries
- Scripting languages such as Python or R
- Data cleaning and validation workflows
- Clinical and molecular data standards
- Project management for partnership-driven work
Practical notes
- The role is based in San Francisco
- Expect regular interaction with external institutions and vendors
- Work is a mix of relationship building and technical data work
- The team cares about doing data work ahead of need rather than in response to it
- Full-time position with focus on the three verticals described above
Project highlights
- Building the dataset foundation for biology training benchmarks
- Creating differentiated data assets from both proprietary and public sources
- Establishing partnerships with biotechs, CROs, and clearinghouses
- Packaging public data into realistic biological analysis tasks
Why you should apply
You will shape how training data is created for AI systems that operate on real-world R&D tasks. The role touches partnerships, strategy, and deep technical work. If you like owning a hard problem from scoping through delivery, and you want your data to matter in biology, this role gives you that scope. The work directly influences what the models learn and how well they perform, and the problems you solve will stay useful for years.