
Staff Software Engineer, Kubernetes Platform
Job description
About the role
Anthropic seeks a Staff Software Engineer to own the reliability, scale, and performance of the Kubernetes control plane that underpins our AI training and inference fleets. You will own, operate, and extend the Kubernetes scheduler to place topology-sensitive ML workloads across thousands of accelerators while ensuring correctness and stability. You will scale the apiserver, etcd, and controllers to operate at orders of magnitude beyond typical limits, identifying bottlenecks before they impact production. You will design, build, and operate core cluster services such as service discovery that every workload depends on, and build custom controllers, operators, and CRDs to automate fleet-specific behaviors. You will partner with research, training, and inference teams to translate workload requirements into platform capabilities, and collaborate with cloud providers on features and escalations. You will participate in on-call rotations, lead incident response, and design processes including postmortems, runbooks, and SLOs to prevent recurrence. Your work will directly determine whether Anthropic can keep reliably and safely training frontier models as our compute footprint continues to grow.
Key facts
What you'll do
- Own, operate, and extend the Kubernetes scheduler for Anthropic's accelerator fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemption.
- Scale the Kubernetes control plane (apiserver, etcd, controller-manager) to support clusters far beyond typical limits, and find the next bottleneck before it finds us.
- Design, build, and operate core cluster services such as service discovery that every workload in the fleet depends on.
- Build and maintain custom controllers, operators, and CRDs to encapsulate fleet-specific logic and automate operational tasks.
- Partner with research, training, and inference teams to understand workload shapes and translate their requirements into platform capabilities.
- Collaborate with cloud providers on required features, escalations, and joint debugging across network, compute, and storage boundaries.
- Participate in on-call rotation, lead incident response, and design processes including postmortems, runbooks, and SLOs that help the team avoid repeating failures.
- Instrument and analyze scheduler and control plane performance at scale to drive decisions on capacity, reliability, and efficiency.
- Contribute to open source Kubernetes projects where appropriate and align internal implementations with upstream best practices.
- Mentor engineers on Kubernetes internals, scheduling concepts, and debugging techniques across the organization.
- Define and implement strategies for upgrading Kubernetes versions safely across a large, diverse fleet.
- Ensure scheduling and control plane behavior meets strict security, compliance, and operational standards.
Requirements
- Significant software engineering experience building and operating production distributed systems.
- Proficiency in at least one systems-appropriate language such as Go, Python, Rust, or C++.
- Deep, hands-on Kubernetes experience well beyond user of, including the scheduler, controllers, apiserver, or operating large multi-tenant clusters.
- Demonstrated ability to debug complex issues across the stack from API behavior down to node and network-level root causes.
- A track record of designing for reliability, correctness, and clear failure semantics in systems other engineers depend on.
- Strong written and verbal communication and comfort building consensus with internal stakeholders.
Nice to have
- Experience with Kubernetes internals or contributions to kube-scheduler, scheduling framework, apiserver, etcd, client-go, controller-runtime, or similar.
- Experience building or operating cluster schedulers or batch systems such as Kueue, Volcano, Slurm, or in-house equivalents.
- Background scaling control planes or coordination systems such as etcd, ZooKeeper, Consul, or large DNS/service-mesh deployments.
- Familiarity with ML infrastructure including GPUs, TPUs, or Trainium; gang scheduling; topology-aware placement; and collective networking such as NCCL.
- Experience with GCP and/or AWS, including GKE/EKS internals and Infrastructure as Code.
- Low-level systems experience such as Linux kernel tuning, cgroups, or eBPF.
- 12+ years of relevant industry experience, including time leading large, ambiguous infrastructure projects.
Practical notes
The compensation range for this role is listed in the compensation section.