Site Reliability Engineer
Job description
About the role
The team depends on this role to own the digital infrastructure that supports Neuro-AI research. Success in this position means compute, registries, and dashboards are reliable and accessible. The role spans compute access, resource visibility, auto-scaling, access management, reproducibility, and process automation.
Platform and DevOps engineers build the systems that run everything else. They manage infrastructure, CI/CD pipelines, observability, and reliability. The work is about automation, scaling, and removing friction for product teams. Systems thinking is the core skill. Platform teams are measured by developer velocity and system reliability. Most companies run on-call rotations, and understanding incident response is part of the role.
Key facts
What you'll do
Resource utilization and cluster health are made visible, so teams can observe patterns and respond proactively.
Compute scales automatically based on demand, so capacity aligns with experimental needs.
Access to resources is controlled so that the right people reach the right systems at the right time.
Deployments and research environments move toward deterministic and reproducible outcomes, reducing variability.
Operational tasks are automated where it increases efficiency, freeing the team for higher-value work.
The current stack of Ansible, Kubernetes, Docker, Tailscale, Python, Grafana, Prometheus, and Talos Linux is used, while remaining open to better tools.
Requirements
You own reliability and availability when the cluster is unhealthy or capacity is tight.
You understand how schedulers, containers, networking, storage, and hardware interact, and you design for graceful degradation.
You value observability, reproducibility, and clear operational boundaries, leaving systems understandable to others.
You pragmatically support experimental research without forcing rigid production constraints, stabilizing when appropriate and allowing controlled exploration when it accelerates discovery.
You work in-person in Emeryville, CA.
Practical notes
Typical interview steps
Platform interviews usually include an infrastructure scenario, a scripting or coding exercise, and operational questions. Candidates may be asked to design a deployment pipeline or debug an outage. Incident experience and an automation mindset are tested. Interviewers often ask about a past outage and how you handled it. Structured post-incident thinking, not heroics, is what they look for.
Good to know
Site reliability engineering centers on keeping distributed systems observable, resilient, and efficient. Teams in research environments often balance strict production standards with the need for fast, iterative experimentation. Modern observability stacks combine metrics, logs, and traces to maintain clarity across complex infrastructure. Container orchestration platforms abstract deployment and scaling while introducing operational complexity that requires deep systems understanding. Networking, storage, and hardware constraints shape how services perform and degrade under load. Toolchains evolve quickly, so adaptability matters more than mastery of any single stack.
Questions to ask
Worth asking in any interview: how the team measures success, who the role works with daily, what the onboarding looks like, and what the company is trying to achieve this year. Asking what past hires did well is a strong final question. Keep the list short and pick the questions that matter most to you.
Career growth
Platform careers grow from engineer to senior, staff, and platform lead roles. Some people move into SRE leadership or cloud architecture. Breadth across networking, storage, and reliability becomes more important at senior levels. Platform careers reward breadth and calm under pressure. Experience automating your own work is the strongest signal for senior roles.
About the company
Astera is a private foundation on a mission to steer science and technology toward an abundant future for all. We believe the coming years will bring an era of unprecedented scientific and technological advancement as exponential progress in AI converges with central advances in other fields to dramatically accelerate innovation.