
Staff Infrastructure Engineer, Cluster Infrastructure
Job description
About the role
Anthropic builds AI systems that are reliable, interpretable, and steerable. Safe, scalable infrastructure determines how quickly we train new models, run safety experiments, and deliver Claude to millions of users. The Cluster Infrastructure team owns the full lifecycle of compute clusters across cloud providers and our own data centers. We create agent-driven automation for provisioning, lifecycle management, and self-healing, ensuring high-bandwidth, secure-by-default connectivity and rapid recovery from faults. This role sets technical direction at a time when compute scale is expanding faster than almost any company. You will define the long-term vision for how clusters are built, operated, and evolved in both cloud and on-prem environments. You will own the design and implementation of systems that abstract complexity while exposing fine-grained control to diverse user needs. You will be responsible for ensuring that infrastructure is not only performant but also observable, recoverable, and aligned with evolving safety and security standards.
Key facts
What you'll do
Define agent-driven strategies for cluster provisioning, updates, and decommissioning across cloud and on-prem environments.
Create automation that ingests new compute capacity on schedule and maintains security-by-default cluster operations.
Establish and evolve practices for operational excellence, including incident response, postmortem culture, and on-call health.
Design secure, high-bandwidth interconnects between clusters while collaborating on physical build-outs with partner teams.
Drive approaches to scalability, homogeneity, and fault tolerance for clusters supporting safety and inference workloads.
Guide long-term compute, data, and infrastructure strategy with cloud providers and internal research teams.
Mentor engineers and support growth within the cluster infrastructure organization over time.
Analyze tradeoffs in designs for rapidly evolving systems that manage global cluster capacity.
Evaluate emerging hardware and software stacks to determine fit for large-scale, safety-critical AI workloads.
Collaborate closely with platform, networking, and security teams to align on standards and shared abstractions.
Implement declarative patterns for infrastructure that enable reproducibility and reduce operational toil over time.
Champion automation that reduces manual intervention while increasing reliability and consistency of cluster operations.
Partner with SRE and research teams to understand capacity requirements and model future infrastructure constraints.
Lead investigations into infrastructure incidents, driving improvements in detection, resilience, and communication.
Requirements
You hold deep expertise in distributed systems, reliability, and cloud platforms such as Kubernetes, Terraform, AWS, GCP, and Azure.
You demonstrate strong proficiency in at least one systems language such as Rust, Go, or Python, along with fluency in infrastructure-as-code using Terraform.
You possess a track record of leading complex, multi-quarter initiatives that span multiple teams and systems.
You show the ability to build alignment across senior stakeholders and communicate effectively at all levels.
You have a history of making sound technical decisions under uncertainty while balancing speed, safety, and maintainability.
You are comfortable working in ambiguous environments where requirements evolve and priorities shift frequently.
You take ownership of outcomes, demonstrating accountability for both successes and failures in the systems you manage.
You are committed to learning and mentoring, helping less experienced engineers grow through code reviews, design discussions, and pair work.
Nice to have
Experience spanning ten or more years as a software engineer in technical lead roles defining team direction.
Operating large-scale compute infrastructure at hyperscale, managing one hundred or more clusters and ten thousand or more nodes.
Depth in Kubernetes internals, cluster provisioning and management systems, and cluster orchestration approaches such as Mesos or Borg-like systems.
Experience with cloud networking, including VPC design, peering, Shared VPC, Transit Gateway, Cloud Interconnect, Cloud NAT, cross-cloud private connectivity, BGP, and DDoS mitigation.
Experience with cluster and host networking, such as CNI, eBPF, NetworkPolicy, multi-NIC, sFlow, and service mesh including Istio, Envoy, and Linkerd with mTLS.
Experience with cluster security, including pod security standards, admission control, RBAC, least-privilege IAM, node and container hardening, and supply-chain image provenance.
Deep experience with infrastructure-as-code tools like Terraform and Atlantis, plus workflow orchestration systems such as Temporal and Argo Workflows.
Practical notes
Full-time engagement in London, UK.
180,000 GBP per year compensation.