
Staff Software Engineer, Node Infra
Job description
About the role
Anthropic is a rapidly growing AI company dedicated to building reliable, interpretable, and steerable artificial intelligence systems that deliver safe and beneficial outcomes for users and society. The Infrastructure organization sits at the core of this mission, defining the compute foundations that enable cutting edge research and responsible scaling of models like Claude. Node Infra owns the complete lifecycle of accelerator capacity across Anthropic, from ingestion and bring-up of new hardware to fleet scale, automated health diagnostics, and repair workflows that keep thousands of GPUs, TPUs, and Trainium devices online. This role focuses on the systems that determine how quickly we can train new models, how reliably we can run safety experiments, and how effectively we can deliver powerful AI capabilities at scale while maintaining the highest standards of reliability and efficiency. The ideal candidate will define technical strategy, drive cross-team initiatives, and establish operational excellence for one of the industry's largest AI compute fleets.
Key facts
What you'll do
Define and own the technical strategy and roadmap for node lifecycle management, covering ingestion, bring-up, health checking, and automated repair of accelerator hardware.
Drive cross-team initiatives to design, provision, and scale AI clusters across multiple cloud providers and accelerator families while optimizing for performance, cost, and reliability.
Design, implement, and operate observability, detection, and remediation systems that automatically identify and isolate unhealthy hardware, reducing mean time between incidents and minimizing stranded capacity.
Own deep infrastructure problems, from capacity management and efficiency to low-level hardware diagnostics, ensuring the hardest challenges are solved either directly or through effective collaboration.
Partner closely with cloud providers and internal research, inference, and product teams to shape long-term compute, data, and infrastructure strategy in support of Anthropic's product and research goals.
Establish and evolve robust operational practices, including incident response, postmortem culture, on-call rotations, and runbooks that keep large-scale AI infrastructure resilient.
Mentor and support the growth of engineers across the node infrastructure domain, providing technical coaching, design reviews, and career development.
Evaluate and integrate emerging technologies and open source tools, translating complex systems design tradeoffs into clear implementation plans for evolving software stacks.
Implement automation that scales with fleet growth, ensuring that node provisioning, repair, and monitoring keep pace with rapid expansion of Anthropic's AI infrastructure.
Champion reliability and efficiency initiatives, using data-driven analysis to drive improvements in fleet utilization, availability, and total cost of ownership.
Collaborate with platform, networking, and kernel teams to address cross-layer issues and contribute to the broader ecosystem through open source contributions where appropriate.
Continuously refine operational processes and tooling to improve developer experience, reduce toil, and increase the safety and predictability of large-scale AI training and inference workloads.
Requirements
Demonstrate deep expertise in distributed systems, reliability engineering, and cloud platforms, with hands-on experience in Kubernetes, infrastructure as code, and major public cloud providers such as AWS, GCP, and Azure.
Show strong proficiency in at least one systems programming language such as Rust, Go, or Python, and be highly skilled in infrastructure as code tools, particularly Terraform, to define and manage complex environments.
Bring hands-on experience with machine learning accelerators including GPUs, TPUs, and Trainium, and understand the implications of different hardware architectures for training and inference workloads.
Have a track record of leading complex, multi-quarter technical initiatives that span multiple teams, systems, and dependencies while delivering reliable results in demanding environments.
Communicate effectively and build alignment across senior stakeholders, translating technical details into clear narratives and decisions for diverse audiences.
Possess strong ownership and problem-solving instincts, with the ability to navigate ambiguity, prioritize effectively, and drive critical infrastructure projects to completion.
Commit to building inclusive, collaborative teams and mentoring engineers at various levels to elevate the overall capability and impact of the node infrastructure organization.
Maintain a strong bias toward action, balancing strategic thinking with execution to ensure that infrastructure investments directly support Anthropic's safety and reliability goals.
Engage with production systems at scale, understanding the realities of operating high-throughput, latency-sensitive infrastructure that powers large language model training and deployment.
Nice to have
Bring 7 or more years of software engineering experience, including time spent as a technical lead defining direction for engineering teams.
Have experience managing large scale compute infrastructure at hyperscale, involving capacity management, efficiency optimization, and lifecycle automation across thousands of nodes.
Show depth in areas such as Kubernetes internals including the scheduler, autoscaler, kubelet, and Karpenter, or in cluster orchestration systems reminiscent of Mesos or Borg-like environments, or in node provisioning pipelines and automation.
Possess low-level systems experience touching the kernel, virtualization, device drivers, firmware, or hardware health and diagnostics daemons that are critical to fleet reliability.
Demonstrate familiarity with high-performance networking technologies such as EFA, RDMA, and InfiniBand as they apply to distributed machine learning workloads.
Have a proven record of production reliability for high-throughput, latency-sensitive systems that must perform consistently under demanding conditions.
Contribute to relevant open source projects in areas such as the Linux kernel, container runtimes, Kubernetes, or other infrastructure software that supports large scale AI workloads.
Show skill in rapidly understanding complex systems designs, evaluating tradeoffs, and keeping track of rapidly evolving software and hardware landscapes across multiple domains.
Practical notes
This role is based in London, UK, and is subject to local hours and travel expectations as defined by company policy.
Travel may be required within the UK and to Anthropic offices and data center locations for collaboration, hardware validation, and incident response.
Visa sponsorship may be available for eligible candidates, subject to role and immigration requirements.
Candidates must meet all listed minimum qualifications and be able to demonstrate relevant experience through past work, projects, or contributions.
The position involves on call responsibilities as part of the operational model for critical infrastructure, and on call rotations will be assigned according to team needs.
Final hiring decisions will consider demonstrated ability to meet the requirements, alignment with Anthropic's mission, and capacity to contribute effectively within a fast-paced, high-scale environment.