Staff+ Site Reliability Engineer, Safeguards ML Infra
Job description
About the role
You are responsible for designing and operating the production infrastructure that directly enables Claude's safety systems to function reliably in real-world conditions. You own the critical backend services that enforce the token generation path and serve as the gatekeeper for every safeguard deployed to production. This position requires you to stand up safeguards for each new model launch, configure them precisely, and verify their behavior under live traffic with strict operational standards. You will serve as the on-call owner for incidents that touch the safeguard stack and hold rollback authority whenever a deployment threatens model integrity or user safety. Every major release flows through your work, and you are expected to convert manual runbooks into automated pipelines that shrink checklist size over time. You will collaborate closely with research and product teams to translate new safety classifiers into robust production services that meet stringent reliability and security criteria.
Key facts
What you'll do
Launch model releases by standing up, configuring, and verifying safeguards for every new model while acting as the safeguards point of contact during release windows.
Own the off-cycle deployment of new safety classifiers as they ship from research, including canarying rollouts, post-deploy validation, and rapid investigation of discrepancies when anomalies appear.
Verify that the correct safeguards are provably live on the intended models across all deployment platforms such as 1P, AWS Bedrock, and GCP Vertex while detecting and eliminating configuration drift.
Automate yourself out of repetitive manual work by turning launch runbooks into tooling, converting hand-built checks into continuous validation, and transforming one-off deploys into repeatable pipelines that reduce ongoing operational load.
Use Claude aggressively to assist in building and refining automation, tooling, and validation systems while paving the way for safe agentic operations in safety-critical environments.
Build and maintain a comprehensive safeguards registry that tracks what is running in production, on which model, on which platform, and documents when and by whom each safeguard was deployed.
Participate in on-call and operational-duty rotations covering service incidents, model provisioning events, and time-sensitive research and safety launches that require immediate response.
Ensure that deployment platforms remain reliable and observable so that safeguards can be released, validated, and rolled back with minimal risk to downstream users and broader ecosystem integrity.
Continuously refine operational processes by converting postmortem action items into durable tooling, automation, and policy changes that prevent recurrence of high-severity issues.
Collaborate with cross-functional teams to align safeguard deployments with product timelines, compliance requirements, and evolving safety standards across all Anthropic customer touchpoints.
Maintain strong communication with engineering, research, and safety teams to ensure that new capabilities are introduced without compromising stability or security guarantees.
Drive improvements in testing, monitoring, and alerting for safeguard infrastructure to increase confidence in releases and reduce manual oversight over time.
Requirements
Have owned production change management at scale, including deploy pipelines, config management systems, canary analysis, and other mechanisms that ensure verified releases under pressure.
Have run high-stakes releases in the past by serving as a launch captain, incident commander, or release owner for systems where a bad deploy has real consequences and remain energized by operating in critical paths.
Have meaningful on-call experience for production systems, including active incident response and postmortem-driven improvements that lead to process and tooling changes.
Have a demonstrated desire to close operational gaps even when no existing playbook exists, including manually sustaining processes until automation and tooling can be built and hardened.
Have hands-on experience deploying and operating services on major cloud platforms such as AWS and GCP at significant scale and reliability.
Be proficient in Python programming, with Rust experience being a valuable addition but not a mandatory requirement for the role.
Possess strong judgment around safety, reliability, and deployment risk in environments where mistakes can affect many users and downstream systems.
Be comfortable working in a fast-moving research and product environment where priorities can shift quickly and operational rigor must keep pace with innovation.
Nice to have
Only items indicated as preferred in the source are included, and no additional preferences have been specified beyond those listed above.
Practical notes
This role requires travel to San Francisco, CA, Seattle, WA, or New York City, NY as indicated in the location section, and remote arrangements are possible with travel accommodations. No specific hours, visa, or application deadline information is provided in the source material.