
Lead Bioinformatician/Engineer
Job description
About the role
Natera is seeking a Lead Bioinformatics Engineer to own pipeline and infrastructure engineering for a research informatics team working on cell-free DNA (cfDNA) screening in Women's and Organ Health. This is an individual contributor role focused on translating complex scientific requirements into robust, automated analysis pathways. You will own our WDL workflows, the transition toward Nextflow, the AWS infrastructure the team runs on, and the tooling around them. You will inherit a working system from the person who built it, and with them still on the team to teach you, ensuring continuity and rapid onboarding. The position is centered on reducing friction for scientists so they can move from raw data to interpreted results with minimal manual intervention. Our workflows leave clear wins available in cost and runtime for whoever takes them on, offering a clear opportunity to make a measurable impact. Getting scientists to the point where they process their own samples end to end, without an engineer in the loop, is a major thrust of this role.
Key facts
What you'll do
Design, maintain, and iteratively improve WDL workflows that power our cfDNA screening analyses in a production environment.
Drive the adoption and standardization of Nextflow workflows alongside existing WDL pipelines to create a cohesive, maintainable workflow portfolio.
Own and refine the AWS infrastructure that hosts our bioinformatics pipelines using Terraform, focusing on reliability, security, and cost-efficiency.
Implement robust testing, monitoring, and alerting frameworks that detect pipeline failures before they impact scientific timelines.
Build and maintain internal tooling that streamlines triage of production issues and reduces repetitive manual work for the analysis team.
Champion the creation of comprehensive documentation and handoff materials to enable scientists to run and debug pipelines independently.
Use coding agents as a routine part of engineering here, integrating them into implementation, code review, investigation, and operational triage workflows.
Review and validate agent-written code to the same rigorous standards as code written manually, ensuring correctness and reproducibility in all outputs.
Collaborate closely with scientific leads and data scientists to translate new methods into production-ready pipelines and support their initial deployment.
Partner with the production organization to ensure smooth handoff of research capabilities, making new features reliable and easy to adopt at scale.
Establish clear operational ownership for all pipeline components, ensuring accountability for performance, uptime, and scalability.
Continuously evaluate new tools and frameworks, running experiments to determine whether they provide tangible benefits over current solutions.
Foster a culture of learning and knowledge sharing within the team, enabling rapid upskilling on emerging bioinformatics technologies and best practices.
Define and track key operational metrics to demonstrate the impact of infrastructure and tooling improvements on scientific throughput.
Requirements
Hold a degree in Bioinformatics, Computer Science, Computational Biology, or a related field, though equivalent depth built through extensive work experience is fully acceptable.
Bring 4+ years of software or pipeline engineering experience, preferably within life sciences, sequencing, or production-adjacent environments where reliability is critical.
Demonstrate proven ownership of a system end to end in production, including deployment, monitoring, and ongoing maintenance beyond initial development.
Show strong Python proficiency with solid software engineering fundamentals, including writing clean, testable, and maintainable code.
Have hands-on experience with at least one workflow system such as Snakemake, CWL, WDL, or Nextflow, and the ability to quickly master new orchestration tools.
Possess real cloud experience, ideally with AWS, including managing infrastructure as code; Terraform is our primary tool but experience with other IaC platforms is acceptable.
Maintain solid practice with Linux command-line environments, containerization technologies, and version control systems like Git.
Have built at least one AI-assisted workflow that remains in active use, and be prepared to discuss what it replaced and how you validated its correctness and performance.
Demonstrate a track record of rapidly learning unfamiliar tools and complex systems, then applying that knowledge effectively to solve ambiguous problems.
Communicate clearly and professionally in written and verbal formats, with the ability to distill technical details for both technical and non-technical stakeholders.
Embrace ownership mindset, driving issues to resolution without unnecessary escalation while knowing when to seek expert consultation.
Approach problems with structured, analytical thinking, balancing speed of delivery with long-term maintainability and operational robustness.
Nice to have
Experience contributing to open source projects relevant to data processing or scientific computing.
Familiarity with healthcare data privacy regulations and compliance considerations relevant to genomic information.
Background working with high-throughput sequencing data or clinical-grade molecular datasets.
Practical notes
This is a full-time position based in the United States with remote work eligibility.
Candidates must be authorized to work in the United States without sponsorship requirements at the time of hiring.