General DiffUSE Job Application
Job description
About the role
This role engages structural biology, machine learning, and open science practice to define how protein dynamics data informs decision pathways. You own the design of analysis strategies that convert heterogeneous structural experiments into robust, reusable knowledge. You own the curation and interpretation of multiconformer ensembles, ensuring that heterogeneity is quantified and communicated clearly. You own the development of representation learning approaches that capture the dynamics of biomolecules beyond static snapshots. You own the partnership with experimental teams to standardize and open existing datasets for community reuse and cross-project validation. You own the translation of raw diffraction data and 3D biomolecular representations into metrics that support downstream modeling and inference. You own the construction of data pipelines that turn experimental outputs into structured, queryable resources for diverse users. You own the communication of methodological choices and uncertainties to both technical and non-technical stakeholders across distributed teams.
Key facts
What you'll do
Design and execute structural data pipelines that ingest diffraction and processing outputs, transforming them into consistent, versioned resources.
Apply macromolecular ensemble metrics to characterize conformational distributions, linking dynamics to functional hypotheses and model confidence.
Implement representation learning methods for protein dynamics, focusing on geometric deep learning applied to biomolecular coordinate and feature spaces.
Develop and maintain data standards workflows around mmCIF, ensuring metadata completeness, provenance tracking, and interoperability with external databases.
Partner with crystallography and imaging facilities to capture raw experimental data, refine processing heuristics, and build scalable generation campaigns.
Construct open datasets and validation benchmarks that align with community norms, enabling external collaborators to contribute back to shared resources.
Lead analysis projects that integrate 3D vision and structural biology with machine learning, emphasizing models trained on raw data rather than preprocessed structures alone.
Coordinate with data producers, platform engineers, and open-science advocates to define interfaces, documentation, and reproducibility practices.
Drive iterative delivery of analytical products, balancing rapid prototyping with rigorous evaluation and clear documentation of assumptions.
Translate complex biophysical and computational concepts into concise narratives that support decision-making across scientific and operational audiences.
Champion the adoption of open data practices, ensuring that curated structural resources remain accessible, interpretable, and extensible.
Support the growth of internal tools and libraries by diagnosing bottlenecks, improving data access patterns, and validating results at scale.
Mentor collaborators on effective data modeling, query design, and visualization techniques to strengthen the broader analytical community.
Continuously monitor emerging methods in structural biology and machine learning, assessing their relevance to existing workflows and future roadmaps.
Requirements
Comfort with diffraction data processing, structural data pipelines, multiconformer and heterogeneity analysis, and data standards such as mmCIF is required for this position.
Experience in macromolecular ensemble metrics, representation learning for protein dynamics, and machine learning on raw experimental data rather than processed structures is mandatory.
Background in 3D vision and geometric deep learning approaches applied to biomolecular problems is a strict eligibility criterion for advancement.
Capacity to design, execute, and document campaigns that generate large structural datasets and associated metadata under reproducible conditions.
Skill partnering with external collaborators to open, annotate, and standardize existing datasets in alignment with community best practices.
Ability to bridge experimental facilities, data producers, and the open-science community through clear interfaces, documentation, and shared standards.
Strong understanding of protein biophysics, structural biology methods, machine learning principles, and scientific data infrastructure is necessary for effective contribution.
Commitment to open science principles, transparent reporting, and collaborative knowledge building across distributed teams and institutional boundaries.
Practical notes
Work occurs in-person at Astera in Emeryville, California, and aligns with the organization's 501(c)(3) structure. Typical interview steps
Data interviews commonly include a SQL or coding exercise, a statistics question, and a case study. Candidates may be asked to design a metric, interpret an experiment, or build a small model. Some companies give a take-home analysis. Expect questions about past projects and the business impact of your work. Interviewers often evaluate how you communicate uncertainty and business impact, not only the math. Bringing a clean write-up of a past analysis to the interview is well received.
Good to know
The role centers on open infrastructure for protein dynamics using structural biology, machine learning, and computational biophysics. Success requires comfort with interdisciplinary boundaries and a bias toward rapid, iterative delivery. Clear written and verbal communication supports collaboration across distributed scientific teams.
Career growth
Data careers grow toward senior analyst, staff data scientist, or data engineering lead. Many professionals specialize in machine learning, analytics, or infrastructure. Cross-functional work with product and engineering teams becomes more important at senior levels. The field changes quickly, so continuous learning is part of the job. Professionals who can translate numbers into decisions tend to advance fastest.
About the company
Astera is a private foundation on a mission to steer science and technology toward an abundant future for all. We believe the coming years will bring an era of unprecedented scientific and technological advancement as exponential progress in AI converges with central advances in other fields to dramatically accelerate innovation.