Staff Cloud Engineer
Job description
About the role
Graphcore is seeking a Staff Cloud Engineer to join the Cloud Platform Team, where you will design, build, and operate cloud services that power AI systems at scale. In this role, you will own the end-to-end lifecycle of cloud infrastructure, ensuring reliability, performance, and scalability for our proprietary AI systems and commercial hardware. You will work hands-on with cloud integration, validation, performance benchmarking, and optimization across a diverse fleet of servers, switches, and storage devices. Collaboration is central to this position, as you will partner closely with Software Platform, Datacentre Operations, and Product Development teams to deliver robust cloud platforms. The ideal candidate thrives in a fast-paced environment and is comfortable solving complex infrastructure problems with creativity and precision. This role demands deep technical ownership and the ability to translate ambiguous requirements into resilient cloud architectures. You will be instrumental in bridging the gap between cutting-edge AI hardware and the cloud software that unlocks its potential.
Key facts
What you'll do
- Run and expand existing OpenStack-based cloud services while architecting and implementing new deployments to meet evolving demands.
- Build and operate end-user services on internal clouds, translating ambiguous user and product requirements into concrete, working systems with measurable outcomes.
- Create automation for gathering and analyzing metrics and observability data, using it to identify trends, diagnose issues, and report findings to stakeholders.
- Coordinate with Datacentre Operations Engineers to ensure AI systems operate at peak performance, troubleshooting issues in private clouds and driving continuous improvements.
- Set up and validate new Graphcore AI hardware using Continuous Deployment and Infrastructure-as-Code practices across internal and external datacentres, ensuring repeatability and reliability.
- Support internal users by providing technical guidance, diagnosing complex issues, and escalating critical findings to Engineering and QA teams for resolution.
- Relay product-related findings from operational work back to engineering and QA teams, helping to close the loop between deployment and product development.
- Contribute to the design and evolution of cloud services, proposing improvements that enhance scalability, resilience, and operational efficiency.
- Implement monitoring and alerting strategies that provide deep visibility into infrastructure health, enabling proactive identification of potential failures.
- Collaborate with cross-functional teams to define standards and best practices for cloud infrastructure, ensuring alignment with organizational goals.
Requirements
- Strong background in cloud infrastructure operations with a proven track record of managing large-scale, production-critical environments.
- Experience with Infrastructure-as-Code deployment methods, using tools and frameworks to automate the provisioning and management of cloud resources.
- Familiarity with observability tools and practices, including metrics, logging, and tracing, to diagnose issues and optimize system performance.
- Knowledge of high-performance networking and storage systems, with an understanding of how these components impact AI workload performance.
- Comfort working with pre-release hardware and software components, thriving in environments where technology is rapidly evolving.
- Prior experience in IT organizations, datacentres, cloud providers, or orchestration and cloud service development, demonstrating exposure to complex operational challenges.
- Ability to work independently and as part of a distributed team, taking ownership of tasks and driving them to completion with minimal supervision.
- Strong problem-solving skills and a methodical approach to debugging, with the patience to investigate issues that span multiple systems and teams.
Practical notes
This is a full-time role based in London, United Kingdom, with no specific travel requirements outlined. The engagement is permanent, and candidates should be prepared to start as soon as feasible. No visa sponsorship details are provided in the source materials, and the standard work hours align with typical full-time expectations in the UK. The role is part of the Cloud Platform Team within the Software Platform organization, emphasizing collaboration and technical ownership. Candidates should be ready to engage with both internal and external datacentres, supporting the deployment and validation of AI infrastructure. The position is well-suited for engineers who enjoy building systems that directly enable cutting-edge AI research and product development.