
Senior Staff+ Software Engineer, Node Infra
Job description
About the role
The Node Infra team at Anthropic is responsible for overseeing the complete lifecycle of accelerator capacity, which includes processes such as ingestion, provisioning, and automated repair of our compute resources. In this role, you will be tasked with designing and managing systems that ensure the operational integrity of GPU, TPU, and Trainium nodes, which are crucial for advancing AI research and model training initiatives.
Key facts
What you'll do
- Develop and define the technical roadmap for managing the lifecycle of nodes, focusing on aspects like deployment, health monitoring, and automated repair processes.
- Spearhead cross-functional projects aimed at scaling AI clusters across various cloud platforms and types of accelerators.
- Create systems that can automatically identify and resolve hardware problems, thereby enhancing the fleet's mean time between incidents and minimizing stranded capacity.
- Design infrastructure solutions and provide technical guidance on complex projects that involve multiple teams.
- Collaborate with research, inference, and product teams to shape the long-term strategy for compute resources.
- Set operational standards for incident management, including on-call rotations and post-incident reviews.
- Mentor and guide engineers to foster their professional development and technical skills.
- Engage in continuous improvement initiatives to enhance system reliability and performance.
- Analyze system performance metrics to identify areas for optimization and enhancement.
- Participate in the development of best practices for infrastructure management and operational excellence.
- Contribute to the documentation of processes and systems to ensure knowledge sharing within the team.
- Stay updated with industry trends and emerging technologies to inform future infrastructure strategies.
Requirements
- Extensive experience in distributed systems, reliability engineering, and cloud environments such as AWS, GCP, or Azure.
- Proficient in at least one systems programming language, including Rust, Go, or Python.
- Familiarity with Terraform for implementing infrastructure as code practices.
- Practical experience with machine learning accelerators, including GPUs, TPUs, or Trainium.
- Demonstrated ability to lead technical initiatives that span multiple quarters and involve collaboration across various teams.
- Excellent communication skills to effectively align with senior stakeholders and team members.
- A Bachelor's degree or a comparable combination of education and relevant work experience.
Nice to have
- Over 12 years of experience in software engineering, particularly in technical leadership roles.
- Experience in managing large-scale compute infrastructure, specifically with over 10,000 nodes.
- In-depth knowledge of Kubernetes internals, cluster orchestration systems, or node provisioning pipelines.
- Understanding of low-level systems, including kernels, virtualization, device drivers, firmware, or hardware diagnostics.
- Familiarity with high-performance networking technologies such as EFA, RDMA, or InfiniBand.
- Proven track record of maintaining production reliability for systems that require high throughput and low latency.
- Contributions to open-source projects, particularly in areas like Linux, Kubernetes, or container runtimes.
Skills & tools
- Programming Languages: Rust, Go, Python
- Infrastructure as Code: Terraform
- Container Orchestration: Kubernetes
- Cloud Platforms: AWS, GCP, Azure
- Accelerators: GPU, TPU, Trainium
- Networking Technologies: EFA, RDMA, InfiniBand
Practical notes
- Visa sponsorship is available, and the process is supported by an immigration attorney.
- The hybrid work policy mandates that employees attend the office at least 25% of the time.
- Anthropic operates as a public benefit corporation, offering benefits such as equity donation matching, generous leave policies, and flexible working hours.
- We encourage candidates to apply even if they do not meet every single qualification listed.
- To avoid recruitment scams, please ensure that all communications come from an @anthropic.com email address.
About the company
Anthropic is an AI safety company that builds reliable, interpretable, and steerable AI systems. Founded in 2021 by Dario and Daniela Amodei, former VP of Research at OpenAI, Anthropic created the Claude family of AI assistants. The company has raised over $13 billion from investors including Google, Salesforce, and Amazon.