
Senior Staff+ Software Engineer, Kubernetes Platform
Job description
About the role
The Kubernetes Platform team manages the control plane for one of the largest AI compute fleets in the industry. You will solve scaling challenges that exceed standard Kubernetes defaults to ensure our infrastructure remains performant and reliable for training frontier AI models.
Key facts
What you'll do
- Develop and extend the Kubernetes scheduler to support accelerator-heavy workloads, including custom plugins for gang scheduling, preemption, and topology awareness.
- Scale the control plane, including etcd, apiserver, and controller-manager, to handle massive increases in object and node counts.
- Build and maintain core cluster services, custom operators, CRDs, and controllers.
- Partner with research and inference teams to translate workload requirements into platform features.
- Lead incident response, participate in on-call rotations, and define SLOs and postmortem processes.
- Coordinate with cloud providers regarding feature requests and technical escalations.
Requirements
- Significant experience building and operating distributed systems in production.
- Proficiency in Go, Python, Rust, or C++.
- Deep, hands-on expertise with Kubernetes internals beyond standard usage, specifically regarding schedulers, controllers, or large-scale multi-tenant clusters.
- Ability to debug complex issues spanning the entire stack, from API behavior to network and node-level root causes.
- Proven track record of designing reliable systems with clear failure semantics.
- Strong communication skills and the ability to build consensus across teams.
- Bachelor degree or equivalent combination of education and experience in a relevant field.
Nice to have
- Contributions to Kubernetes internals like client-go, controller-runtime, or the scheduling framework.
- Experience with batch systems or cluster schedulers such as Volcano, Kueue, or Slurm.
- Background in scaling coordination systems like ZooKeeper, Consul, or etcd.
- Familiarity with ML infrastructure, including GPUs, TPUs, Trainium, NCCL, and topology-aware placement.
- Experience with GCP, AWS, EKS, GKE, and Infrastructure as Code.
- Low-level systems knowledge, including eBPF, cgroups, or Linux kernel tuning.
- 12+ years of industry experience, including leading large-scale infrastructure projects.
Skills & tools
- Kubernetes (Scheduler, Control Plane, CRDs)
- Go, Python, Rust, C++
- Distributed Systems
- GCP, AWS
- ML Infrastructure (GPUs, TPUs)
Practical notes
- Visa sponsorship is available and we retain immigration counsel to assist with the process.
- We operate under a location-based hybrid policy requiring at least 25% time in the office.
- We encourage applications even if you do not meet every listed qualification.
- Please review our candidate AI usage policy on our website before applying.
- Beware of recruiting scams; only communicate via @anthropic.com email addresses.
About the company
Anthropic is an American artificial intelligence public benefit corporation headquartered in San Francisco, California. The company was founded in 2021 with the goal of promoting AI safety. Its flagship product is Claude, a series of large language models. Anthropic was established by former members of OpenAI, including siblings Daniela Amodei and Dario Amodei. Daniela Amodei serves as president and Dario Amodei serves as chief executive officer. The company is privately held but plans to go public. It had an estimated valuation of $965 billion in May 2026, making it the most valuable pure-play AI company in the world. Anthropic focuses on building reliable, interpretable, and steerable AI systems. The company conducts research across multiple areas including interpretability, alignment, and societal impacts. Its approach emphasizes safety at the frontier of AI capabilities. The public benefit corporation structure allows the company to prioritize long-term safety outcomes alongside commercial objectives. Anthropic works with governments, civil society, and industry partners to promote responsible AI development. The company publishes research papers and shares safety insights with the broader AI community. Its team includes researchers, engineers, and policy experts working on technical and governance challenges. Anthropic aims to ensure that advanced AI systems benefit humanity.