
Staff+ Software Engineer, Node Infra
Job description
About the role
Anthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. You will own the technical strategy and roadmap for node lifecycle management, including ingestion, bring-up, health checking, and automated repair of accelerator hardware. You will drive cross-team initiatives to build and scale AI clusters across multiple cloud providers and accelerator families, ensuring that the hardest problems are solved either directly by you or through coordination with senior engineers. You will design and operate the systems that detect, isolate, and remediate unhealthy hardware automatically, driving up fleet MTBF and minimizing stranded capacity across Anthropic's large-scale infrastructure. You will work closely with cloud providers and internal research, inference, and product teams to shape long-term compute, data, and infrastructure strategy in support of safe and scalable AI deployment. You will establish and evolve operational excellence practices such as incident response, postmortem culture, and on-call regimes while mentoring and coaching engineers to help the Node Infra organization grow and thrive.
Key facts
What you'll do
- Define and own the end-to-end technical strategy for node lifecycle management, covering hardware ingestion, system bring-up, continuous health validation, and automated remediation workflows.
- Architect and operate scalable infrastructure that spans multiple cloud providers and accelerator families, aligning capacity planning with product and research demand.
- Design, implement, and maintain automated detection, isolation, and repair systems for unhealthy hardware to maximize fleet availability and reduce manual intervention.
- Lead complex, multi-quarter initiatives that cut across infrastructure, tooling, and product teams to deliver reliable and efficient AI compute clusters.
- Partner closely with cloud providers, network teams, and internal research, inference, and product groups to influence long-term compute, data path, and infrastructure roadmaps.
- Build and sustain strong operational practices, including incident response, observability, postmortem processes, and on-call rotations for critical node infrastructure.
- Provide technical mentorship and coaching to engineers across the node lifecycle stack, fostering growth, clarity, and ownership within the team.
- Translate ambiguous product and research requirements into robust infrastructure designs that balance performance, reliability, cost, and deployability.
- Contribute to open-source projects and industry standards where relevant, ensuring Anthropic's infrastructure work benefits the broader ecosystem.
- Continuously evaluate new hardware, software, and tooling to determine fit for Anthropic's scale, iterating quickly where experiments show clear advantages.
Requirements
- Demonstrate deep expertise in distributed systems, reliability engineering, and cloud platforms such as Kubernetes, Infrastructure as Code, and major public cloud providers.
- Show strong proficiency in at least one systems-level programming language such as Rust, Go, or Python, along with hands-on experience using Terraform or similar IaC tools.
- Bring hands-on experience with machine learning accelerators including GPUs, TPUs, or Trainium, including driver, firmware, and hardware-level interactions where applicable.
- Provide a track record of leading complex, multi-quarter technical initiatives that span multiple teams, systems, and dependencies without close supervision.
- Illustrate the ability to build alignment across senior stakeholders and communicate technical tradeoffs clearly and persuasively at all organizational levels.
- Evidence of ownership in production environments where high throughput, low latency, and strict reliability are non-negotiable for demanding workloads.
- Prior experience managing large-scale compute infrastructure at hyperscale, including capacity planning, efficiency optimization, and cost-aware decision-making.
- Comfort with low-level systems concepts such as kernel behavior, virtualization mechanisms, device drivers, firmware interfaces, or hardware health and diagnostic daemons.
Nice to have
- Experience with Kubernetes internals including the scheduler, autoscaling components, kubelet, and cluster autoscaling solutions such as Karpenter.
- Familiarity with cluster orchestration systems that resemble Borg or similar large-scale schedulers used in hyperscale environments.
- Hands-on background in node provisioning pipelines, from bare metal or VM images to fully configured and validated accelerator nodes.
- Contributions to relevant open-source projects such as the Kubernetes project, Linux kernel, container runtimes, or related infrastructure software.
- Demonstrated skill in understanding and communicating systems design tradeoffs in fast-moving, high-growth engineering organizations.
- Exposure to high-performance networking technologies like EFA, RDMA, and InfiniBand in the context of distributed machine learning workloads.
Practical notes
This role is based in the United States in San Francisco, California; New York City, New York; or Seattle, Washington. Employment is full-time. Candidates must be authorized to work in the country where they apply. Relocation support may be available for qualified candidates. Compensation details are provided in transparent ranges to support pay transparency.