Informatics Engineer
Job description
About the role
Prime Medicine is a leading biotechnology company dedicated to creating and delivering the next generation of gene editing therapies to patients. The Company is leveraging its proprietary Prime Editing platform, a versatile, precise and efficient gene editing technology, to develop a new class of differentiated, one-time, potentially curative genetic therapies. Designed to make only the right edit at the right position within a gene while minimizing unwanted DNA modifications, Prime Editors have the potential to repair almost all types of genetic mutations and work in many different tissues, organs and cell types. The Informatics Engineer will own the end-to-end delivery of the data and compute backbone that supports these therapeutic ambitions. You will be responsible for designing and operating the critical infrastructure that connects experimental science with production-grade software. This role requires deep collaboration with both computational biology and cloud infrastructure teams to ensure data integrity and system reliability. Your engineering decisions will directly influence how scientific teams interact with complex genomic datasets on a daily basis. You will own the design, implementation, and evolution of the platform that powers our core data workflows.
Key facts
What you'll do
Design and build the data platform connecting NGS instruments, laboratory informatics systems (Benchling), and AWS cloud compute; covering automated ingestion, provenance tracking, scalable storage, and observability.
Build production-grade APIs, SDKs, and internal tools that bring genomic data and analytical capabilities to scientists across the organization.
Build the orchestration layer for the automated, event-triggered execution of our scientific pipelines. The pipelines themselves (amplicon-seq, off-target analysis, and others) are co-developed with our computational biology team; you'll own how they run, scale, and integrate.
Ship AI-powered and agentic capabilities such as RAG over internal scientific data, agentic workflows with human-in-the-loop review, and internal copilots that streamline routine workflows across the organization connecting both science and business needs.
Build integrations across the scientific tool chain so data moves reliably between ELNs, LIMS, instrument software, and cloud compute.
Translate scientific requirements into reliable, maintainable software, helping bring research prototypes into production-grade systems.
Partner with computational biologists, lab scientists, and our cloud-support team to continuously improve platform performance, cost, and reliability.
Implement robust data models and storage strategies to handle the complexity and scale of genomic datasets within a regulated environment.
Ensure platform security, compliance, and auditability as data moves from experimental instruments to long-term storage and analysis.
Champion best practices in software engineering, including testing, monitoring, logging, and documentation for scientific data workflows.
Work closely with product teams to scope and deliver data-centric features that unblock scientific discovery and development timelines.
Continuously evaluate new technologies and methodologies to optimize the efficiency and cost-effectiveness of the data platform.
Requirements
5+ years of full-time engineering experience in production environments (3+ for those with an MS or PhD).
Strong Python programming skills with demonstrated experience in scientific or data-intensive applications.
Production AWS experience with depth in event-driven, cloud-native architectures, including AWS Lambda, EventBridge (or comparable event-routing infrastructure), and Infrastructure as Code with Terraform or CDK.
Docker / containerization, Git-based workflows, CI/CD, code review, and testing code in DEV/TEST environments before deploying to production.
A track record of building integrations and automations, preferably involving scientific or regulated data streams.
Experience with data pipeline frameworks, workflow orchestration tools, and database systems relevant to genomic data.
Strong understanding of data modeling, schema design, and query patterns for analytical workloads.
Commitment to operational excellence, including monitoring, alerting, and incident response for data platforms.
Nice to have
Experience in life-science related industry involving scientific or regulated data, such as bioinformatics platforms, electronic lab notebooks, or laboratory information systems.
Familiarity with genomic data formats, analysis workflows, and data standards commonly used in research and clinical settings.
Experience with RAG, agentic workflows, or AI tooling for internal productivity and decision support in scientific contexts.
Knowledge of compliance frameworks relevant to regulated data in life sciences.
Practical notes
Engagement is Full-time.
The role is based in Cambridge, MA.
Candidates must be eligible to work in the United States without sponsorship for this position.