Head of Platform Infrastructure
AnyscaleUSAFull Time3w ago
Job description
About the role
Anyscale is seeking an experienced engineering leader to oversee our Infrastructure, SRE, and Enterprise Governance Engineering teams. In this role, you will be responsible for defining and executing the technical vision for the platform infrastructure that supports scalable distributed AI applications. You will lead a talented team of engineers, guiding their efforts to build reliable, scalable, and efficient systems that enable developers and customers to deploy Ray-based applications seamlessly in cloud environments. Your leadership will be crucial in driving strategic initiatives, ensuring operational excellence, and fostering a culture of innovation and high performance within the organization.
Key facts
What you'll do
- Lead and shape the technical strategy for core infrastructure components including cluster launcher, cloud provider integrations (AWS, GCP, Azure), Kubernetes support, cluster autoscaling, and the control and data planes.
- Oversee the development, deployment, and reliability of the billing stack, production database, and related infrastructure that support the platform's operational needs.
- Manage and grow a high-performing engineering team focused on infrastructure, site reliability engineering, and enterprise governance, including recruiting, mentoring, and performance management.
- Collaborate closely with customers and field engineering teams to understand their challenges, gather feedback, and ensure their success with the platform.
- Drive the execution of projects aimed at improving scalability, performance, and reliability of the platform infrastructure, ensuring timely delivery and quality standards.
- Establish and promote best practices for distributed systems, cloud integrations, and infrastructure automation to ensure robustness and maintainability.
- Prioritize initiatives based on business needs, technical feasibility, and operational impact, making data-driven decisions to guide the team's efforts.
- Develop and implement monitoring, alerting, and incident response strategies to maintain high system availability and reliability.
- Foster a culture of continuous improvement, ownership, and innovation within the engineering teams, encouraging knowledge sharing and technical excellence.
- Work cross-functionally with product management, security, and other engineering teams to align infrastructure development with overall company goals and compliance requirements.
- Ensure adherence to security standards and manage risks related to cloud infrastructure, including compliance with relevant regulations and policies.
- Represent the platform infrastructure team in executive meetings, providing updates on progress, challenges, and strategic plans.
Requirements
- Proven leadership experience managing engineering teams in infrastructure, SRE, or related areas, with a track record of building high-performing teams.
- Deep technical expertise in distributed systems, cloud infrastructure, Kubernetes, and virtual machines.
- Extensive experience working with cloud providers such as AWS, GCP, and Azure, including their APIs, services, and best practices.
- Demonstrated success in scaling teams and systems quickly while maintaining a strong engineering culture.
- Strong project management skills, with the ability to deliver complex projects on schedule and within scope.
- Excellent communication skills, capable of articulating technical strategies to both technical and non-technical stakeholders.
- Ability to motivate and mentor engineers, handle performance management issues, and foster professional development.
- Experience with cloud automation tools, infrastructure as code, and monitoring/alerting systems.
- Knowledge of enterprise governance, security standards, and compliance considerations.
- Ability to handle high-pressure situations and prioritize effectively to meet organizational goals.
- A proactive mindset with a sense of urgency and a focus on achieving results.
Nice to have
- Prior experience working on distributed AI or machine learning platforms, especially with Ray or similar frameworks.
- Familiarity with enterprise governance, compliance, and security standards relevant to cloud infrastructure.
- Experience supporting large-scale cloud deployments and managing billing systems.
- Knowledge of supporting and managing production databases and data infrastructure.
- Background in supporting or developing cloud-native applications and services.
Skills & tools
- Distributed systems architecture and design
- Kubernetes and container orchestration
- Cloud platforms: AWS, Azure, GCP
- Infrastructure automation and management tools (e.g., Terraform, CloudFormation)
- Monitoring and alerting systems (e.g., Prometheus, Grafana)
- Cloud provider APIs and SDKs
- Programming skills relevant to infrastructure and automation (e.g., Python, Go)
- Security best practices for cloud infrastructure
- CI/CD pipelines and DevOps practices
Practical notes
- This is an on-site role based in San Francisco.
- Candidates must have legal authorization to work in the United States; Anyscale is an E-Verify company.
- The role involves managing sensitive information subject to ITAR regulations.
- Anyscale is an equal opportunity employer committed to diversity and inclusion.
- The company has raised over $250 million and is backed by Andreessen Horowitz, NEA, and Addition.
- The position requires a candidate with a strong technical background, leadership skills, and the ability to work effectively across teams and with customers.