
Staff Software Engineer, Inference
Job description
About the role
The Inference team builds the infrastructure required to serve Claude to millions of global users. You will manage the entire stack, from traffic routing to fleet orchestration, across diverse AI hardware and cloud environments. In this role, you will own the design and execution of critical infrastructure components that directly impact the performance and reliability of Claude's serving stack. You will collaborate closely with hardware engineers, researchers, and platform teams to ensure that the serving infrastructure keeps pace with rapid model development. Your work will involve making high-stakes architectural decisions that balance scalability, efficiency, and operational simplicity. You will be responsible for translating complex production requirements into robust, maintainable software systems. This position is for an engineer who thrives on solving ambiguous, large-scale problems that span software, systems, and machine learning.
Key facts
What you'll do
- Architect and maintain distributed systems that power Claude for a massive user base, ensuring high availability and low latency.
- Create flexible infrastructure that adjusts in real time to shifting production demands, optimizing resource utilization and cost.
- Build traffic management, load balancing, and request routing systems for thousands of accelerators, handling complex routing logic.
- Optimize compute efficiency through automated scaling of research and production workloads, reducing waste and improving throughput.
- Manage deployment pipelines for releasing new models to production, implementing robust testing and rollout strategies.
- Provide high-performance infrastructure to support next-generation model development, enabling faster experimentation and iteration.
- Integrate new AI accelerator platforms and support novel model architectures, ensuring compatibility and performance.
- Implement monitoring and observability solutions to gain deep insights into system behavior and quickly identify issues.
- Drive standardization across services and components to simplify operations and improve developer experience.
- Partner with security and reliability teams to ensure infrastructure meets stringent compliance and safety requirements.
- Troubleshoot intricate production incidents, coordinating with cross-functional teams to resolve issues swiftly.
- Contribute to the design of disaster recovery and failover mechanisms to protect against service disruptions.
- Mentor junior engineers by providing guidance on best practices for distributed systems and infrastructure as code.
- Evaluate and prototype new technologies, assessing their potential impact on the serving stack and operational workflows.
Requirements
- Proficiency in Python or Rust, with the ability to write efficient, idiomatic code in at least one of these languages.
- Experience building and running distributed systems in production environments, with a deep understanding of their complexities.
- Practical knowledge of containerized infrastructure like Kubernetes and at least one major cloud provider (AWS, GCP, or Azure).
- Bachelor degree or equivalent combination of training, education, and experience in a relevant field.
- Ability to thrive in environments where technical output drives business and research goals, adapting quickly to changing priorities.
- Interest in the societal implications of AI development, considering how infrastructure choices affect model deployment and access.
- Willingness to take on tasks outside of your core responsibilities, demonstrating ownership and a proactive mindset.
- Strong problem-solving skills, with the patience to debug difficult issues in complex, distributed systems.
- Excellent communication skills, enabling effective collaboration with engineers, product managers, and researchers.
Nice to have
- Extensive experience with large-scale, high-performance distributed systems, having operated services at massive scale.
- Background in deploying machine learning systems at scale, with a track record of successful production deployments.
- Expertise in building request routing, load balancing, or traffic management systems, particularly for AI workloads.
- Familiarity with LLM inference optimization, including caching and batching strategies, to maximize throughput and reduce latency.
- Deep operational experience with Kubernetes and cloud infrastructure, including networking, storage, and security.
- Experience with AI hardware such as TPUs, GPUs, or emerging accelerator platforms, and understanding their implications for software design.
- Knowledge of infrastructure as code tools and automation frameworks for managing cloud resources.
- Experience with observability tools and practices, including logging, metrics, and tracing, to maintain system health.
- Familiarity with security best practices for cloud-native applications and distributed systems.
- Understanding of cost optimization strategies for large-scale cloud infrastructure and AI workloads.
Practical notes
- Hybrid work policy: Staff are expected to be in the office at least 25% of the time.
- Visa sponsorship: Available. The company retains immigration counsel to assist with the process.
- Applications are reviewed on a rolling basis with no fixed deadline.
- Candidates are encouraged to apply even if they do not meet every listed qualification.
- Verify all communications originate from an @anthropic.com email address to avoid recruitment scams.