Platform Engineering Lead | AI Infra
Job description
About the role
You will own the compute, data, AI, and security infrastructure that powers every sequencing run in our operation. You will orchestrate the complex interplay between lab robots, sequencers, and on-prem Linux servers while ensuring high quality data ingestion. You will design and build the cloud infrastructure and data pipelines that reliably house, process, and deliver terabytes of sequencing data to our customers. You will architect systems to orchestrate millions of bioinformatics tasks using queues and workflow engines for maximum throughput. You will spearhead the development of AI infrastructure and internal tooling to enable autonomous agentic systems across the business. You will partner closely with scientists and engineers to translate ambitious product goals into reliable, scalable production services. You will drive incident response and performance optimization for production platforms where reliability is non-negotiable.
Key facts
What you'll do
Design and maintain the core services that orchestrate lab devices and the data ingestion pipeline for high-throughput sequencing operations.
Architect cloud infrastructure and data pipelines capable of ingesting, processing, and serving terabytes of sequencing data with low latency.
Design infrastructure to schedule and orchestrate millions of bioinformatics tasks using robust queues and workflow management systems.
Build AI infrastructure and internal tooling to support agentic systems that can operate with minimal human supervision.
Develop Quality Scientist Agents that monitor operations end-to-end, surface anomalies, and trigger interventions or escalations when quality or reliability issues arise.
Develop Logistics Agents that coordinate drivers and carriers globally to ensure samples and packages move efficiently through the pipeline.
Develop Bioinformatics Coding Agents that perform long-running, adaptive analysis on diverse sample types with shifting data distributions.
Implement strong REST API best practices and service-oriented architectures to ensure interoperability across services and teams.
Apply rigorous SQL and data modeling skills in relational databases such as postgres to manage complex sequencing metadata and results.
Write clean, scalable Python backend services using frameworks such as Flask, Django, or FastAPI to support internal and external consumers.
Operate production systems on AWS using containerized task orchestration with Lambda, ECS, Batch, and Step Functions for resilient workflows.
Utilize infrastructure-as-code tools such as Terraform or AWS CDK to provision and govern cloud resources safely and reproducibly.
Build and maintain distributed systems powered by task queues such as Celery and SQS to handle bursty and long-running workloads.
Administer Linux servers using systemd services, basic networking knowledge, and application deployment tools such as ansible for automation.
Comfortably operate production systems by leading incident response, debugging complex issues, and driving performance and reliability improvements.
Champion the adoption of modern development practices as primarily developing with coding agents across design, implementation, test, and review.
Foster a culture of learning by collaborating with scientific and engineering partners to understand DNA sequencing bioinformatics algorithms and workflows.
Promote clear communication, mentorship, and knowledge sharing to help team members grow new skills and advance their careers.
Push innovative solutions to production in a fast-paced environment where agility and ownership are highly valued.
Requirements
BS in Computer Science, Computer Engineering, Math, Physics or technical subject.
7+ years of industry experience or equivalent hands-on technical background.
Prior experience leading projects and mentoring engineers through technical and people leadership.
Demonstrated experience building and deploying LLM/agent systems including tooling, evaluation, safety, and permissions design.
Extensive background developing with coding agents across all parts of the software lifecycle including design, implementation, test, and review.
Strong expertise with REST API best practices such as OpenAPI and service-oriented architectures for distributed systems.
Strong mastery of SQL and data modeling in relational databases especially postgres for complex datasets.
Strong proficiency in Python backend services using frameworks like Flask, Django, or FastAPI in production settings.
Deep experience working with AWS for containerized task orchestration using Lambda, ECS, Batch, and Step Functions at scale.
Hands-on experience with infrastructure-as-code tools such as Terraform or AWS CDK for cloud resource management.
Proven track record building and operating distributed systems leveraging task queues such as Celery and SQS for reliable concurrency.
Experience managing Linux servers with systemd services, basic networking concepts, and deployment tools like ansible for automation.
Strong desire to learn about DNA sequencing bioinformatics algorithms and workflows to drive better engineering decisions.
Commitment to an inclusive environment where diverse and creative perspectives are welcomed and encouraged.
Ability to thrive in a fast-paced, high-agency setting with tight-knit collaboration and ambitious goals.
Nice to have
Experience contributing to open source projects relevant to data pipelines and infrastructure.
Familiarity with bioinformatics workflows and standards in genomics research.
Practical notes
This role is full_time based in San Francisco.
Plasmidsaurus is an equal opportunity employer and does not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. If you need additional accommodations to feel comfortable during your interview process, please contact careers@plasmidsaurus.com.