
Senior Software Engineer I/II, Back-end/Data, Robotics
Job description
About the role
You design and operate the back-end systems that power autonomous laboratories. Your work turns experimental plans into executable workflows, instrument signals into governed data, and physical factories into predictable digital twins. You own the core services that schedule work, manage robots, and ensure scientists can trust measurements and act on them with confidence. This role combines infrastructure, data engineering, and scheduling logic to support high-throughput, regulated research environments. You will be responsible for the reliability and correctness of the systems that control physical scientific hardware. The position requires a deep understanding of how software guarantees translate into trustworthy laboratory outcomes. You will partner closely with both engineering and scientific teams to solve complex problems at the intersection of software and biology.
Key facts
What you'll do
- Architect durable scheduling workflows that allocate instruments, compute, and experiments across autonomous labs using long-running, durable execution units.
- Build and maintain digital twin simulation layers that mirror physical factory behavior, enabling teams to test scenarios before resource commitment.
- Implement technical data management with Flyte pipelines, an S3 lakehouse, and PostgreSQL models to transform raw instrument output into reliable, queryable assets.
- Create event-driven data pipelines that deliver clean datasets to scientists and machine learning models, supporting both batch and streaming needs.
- Develop capacity planning services that provide forecasts and throughput simulations for operations and customer commitments.
- Operate production services on cloud infrastructure, ensuring reliability, scalability, and performance as laboratory volume grows.
- Drive services from initial design through deployment, monitoring, and iterative improvement with clear ownership and accountability.
- Translate complex technical tradeoffs into clear narratives for engineers, scientists, and business stakeholders to align priorities and constraints.
- Define and enforce interfaces between robotic controllers, scheduling systems, and data stores to maintain consistency across the lab ecosystem.
- Analyze production incidents and reliability metrics to drive improvements in system resilience and observability.
- Collaborate with data scientists to ensure analytics and modeling workloads have performant and secure access to laboratory data.
- Contribute to open source tooling where appropriate to maintain best practices and leverage community innovation.
- Optimize storage and compute costs in object storage and data lakehouse environments without sacrificing data integrity or access patterns.
- Support the design of experiments through programmatic interfaces that abstract the complexity of the underlying robotics fleet.
- Ensure all services comply with data governance policies and regulatory requirements relevant to life sciences research.
Requirements
- Bring 4-8 years of back-end engineering experience, with production Python services at scale using FastAPI.
- Have hands-on experience with data engineering pipelines coordinated by workflow orchestrators such as Flyte, Temporal, or similar systems.
- Model and work with data across multiple stores, including SQL in PostgreSQL, object storage on S3, and data lakehouse architectures.
- Deliver cloud-native solutions using containers, infrastructure as code, and CI/CD pipelines, with proven experience on AWS container services and infrastructure tooling.
- Own end-to-end lifecycles for services, including design decisions, implementation, reliability monitoring, and measured iteration based on feedback.
- Communicate plainly about technical tradeoffs with cross-functional stakeholders, including engineers, scientists, and business partners.
- Write comprehensive tests and maintain high standards for code quality, documentation, and operational runbooks.
- Demonstrate the ability to debug complex distributed systems issues using logs, metrics, and traces.
- Be comfortable working in a fast-paced environment where requirements evolve based on scientific discovery and customer needs.
Nice to have
- Experience in regulated industries where systems engineering rigor is essential for safety or quality, such as aerospace, manufacturing, or healthcare and life sciences.
- Domain exposure in robotics, lab automation, LIMS, or manufacturing/MES environments.
- Background in scheduling, resource allocation, or constraint and optimization problems.
- Experience enabling data science with distributed compute frameworks, experiment tracking systems, or model serving infrastructure.
- Familiarity with event-driven messaging platforms such as NATS or MQTT for high-volume telemetry and asynchronous coordination.
Skills & tools
Python; FastAPI; Flyte; Temporal; PostgreSQL; S3; AWS EKS; Docker; ECR; Terraform; GitHub Actions; NATS; MQTT.
Practical notes
Full-time role located in Cambridge, Massachusetts. Candidates must be authorized to work in the United States without sponsorship for this position. International candidates will receive region-tailored benefits if eligible under local regulations. Travel is not required for this role. The position reports to the engineering organization in Cambridge and participates in on-call rotations as needed.