Staff Software Engineer, AI Reliability Engineering
Job description
About the role
This role focuses on ensuring Claude remains dependable for every user who relies on it. You will own the reliability of critical AI serving paths that span from the SDK to the network and into the accelerators that power inference. The position requires you to partner deeply with infrastructure, product, and safety teams to remove systemic risk from the stack. You will design observability that exposes subtle failure modes before they impact users. Your work will directly influence how resilient Anthropic's most important services are during incidents and long term roadmap initiatives. You will be expected to operate in ambiguous situations and make sound decisions with incomplete information. Ultimately, you will help define what reliability means for a modern AI system at scale.
Key facts
What you'll do
Investigate emergent failure modes across the token path and translate them into durable architectural improvements.
Define and evolve Service Level Objectives that balance latency, availability, and development velocity for large language model services.
Construct monitoring and observability systems that provide signal-rich insights across the entire serving stack.
Lead the design of high-availability infrastructure that spans multiple regions and cloud providers.
Own incident response playbooks for critical AI services and drive rapid recovery and methodical postmortems.
Implement safeguards that ensure reliability improvements keep pace with the evolution of model capabilities.
Collaborate with partner teams to embed reliability practices into their design and delivery workflows.
Evaluate and adopt new observability tools and frameworks specific to AI and machine learning workloads.
Perform readiness reviews for major deployments, identifying weak spots in distributed systems and model serving paths.
Champion resilience testing practices such as chaos engineering to validate system behavior under stress.
Build trusted relationships with cross-functional teams to act as a cohesive unit rather than a detached advisory group.
Represent the reliability perspective in product and infrastructure decisions to prevent future class issues.
Analyze large-scale model serving and training infrastructure to surface risks before they escalate.
Contribute to open-source infrastructure and ML tooling that strengthens the broader reliability community.
Requirements
Bachelor's degree or an equivalent combination of education, training, and/or experience is required.
A field of study relevant to the role as demonstrated through coursework, training, or professional experience is mandatory.
Years of experience required will correlate with the internal job level requirements for the position.
You must have strong distributed systems, infrastructure, or reliability backgrounds.
You should be a reliability-minded software engineer or Site Reliability Engineer focused on large scale systems.
You must be comfortable jumping into unfamiliar systems during an incident and helping drive resolution without needing full expertise.
Holistic thinking is required to understand how systems compose and where the seams between them are located.
You need to build lasting relationships across teams and be welcomed as a teammate rather than an outsider with opinions.
Ownership over user outcomes is essential, even for systems you do not directly own.
Excellent communication and collaboration skills are required due to the cross-company engagement model.
Experience operating large-scale model serving or training infrastructure involving more than 1000 GPUs is expected.
Familiarity with one or more ML hardware accelerators such as GPUs, TPUs, or Trainium is necessary.
Understanding of ML-specific networking optimizations like RDMA and InfiniBand is required.
Expertise in AI-specific observability tools and frameworks is a strict requirement.
Experience with chaos engineering and systematic resilience testing is expected.
Contributions to open-source infrastructure or ML tooling are part of the hard bar expectations.
Nice to have
Preferred items are not specified in the source; there are no additional nice to have qualifications listed.
Practical notes
The role is full_time.
No specific hours, travel, visa, or deadlines are stated in the source text beyond the general expectation to apply even if you do not meet every qualification.
Visa sponsorship is available, although not guaranteed for every role or candidate.
Encouragement is provided to apply regardless of perceived qualification gaps, especially for underrepresented groups.