Senior Cloud Software Engineer
Job description
About the role
The role manages cloud infrastructure to support AI workloads across hardware development. The position focuses on deploying scalable systems within a distributed engineering organization. You will own the design and implementation of infrastructure that directly enables machine learning innovation. Collaboration with distributed teams is central to ensuring reliable and performant cloud operations. The work involves translating product requirements into robust cloud-based solutions. You will play a key part in automating and securing the environment for AI workloads. Success in this role requires clear communication to explain complex technical decisions effectively.
Key facts
What you'll do
Design and implement infrastructure components that provide the foundation for large-scale AI model training and inference workflows.
Automate deployment pipelines to enable rapid and reliable release cycles for cloud-based AI services managed by the team.
Operate production cloud environments with a focus on performance, reliability, and efficient capacity planning derived from observability data.
Respond to production incidents by analyzing observability data and collaborating with peers to restore service stability quickly.
Apply strong networking, security, and identity management principles to safeguard AI workloads and sensitive data flows at every layer.
Partner with hardware and software teams to align cloud infrastructure strategies with the evolving needs of AI processor development.
Utilize infrastructure-as-code practices to define, provision, and manage cloud resources in a consistent and version-controlled manner.
Champion automation initiatives to reduce manual operational tasks and improve the resilience of the overall cloud platform.
Analyze system metrics and logs to identify bottlenecks and guide long-term capacity planning for critical production services.
Contribute to code reviews and technical documentation to ensure high standards are maintained across the distributed engineering organization.
Requirements
Strong understanding of networking, security, and identity management in cloud environments to protect AI workloads and data flows effectively.
Demonstrated experience with cloud platforms and infrastructure-as-code tools is essential for deploying scalable systems.
Proven ability to work within a distributed engineering organization where communication and clarity are as valued as technical skill.
Solid experience in building, monitoring, and operating production-grade cloud infrastructure is required for this position.
The ability to explain technical decisions in plain words is a necessary trait to succeed in this role.
Experience with automation, observability, and monitoring tools is required to manage performance and reliability.
You must be comfortable working with AI workloads and the unique demands of training and inference in cloud environments.
Full-time engagement is required, and the role involves significant responsibility for critical cloud systems.
Practical notes
Typical interview steps
Hiring for engineering roles usually starts with a recruiter screen, followed by one or two technical rounds. Candidates often solve a coding problem, discuss past projects, and answer system design questions. Some loops include a take-home task. Final rounds typically cover team fit and give candidates a chance to ask questions. Interviewers look for how you break down an unfamiliar problem, not just whether you reach the answer. Practicing a few problems aloud and reviewing your own past projects are the best preparation.
Good to know
Work in this field centers on cloud infrastructure for AI workloads, often using containers and infrastructure-as-code tools. Teams rely on strong communication in distributed collaboration. Observability, networking, and automation guide operations in production environments.
Questions to ask
Worth asking in any interview: how the team measures success, who the role works with daily, what the onboarding looks like, and what the company is trying to achieve this year. Asking what past hires did well is a strong final question. Keep the list short and pick the questions that matter most to you.
Career growth
Engineering careers usually progress from individual contributor to senior, staff, and principal levels. Some engineers move into management and lead teams of five to twenty people. Others stay on the technical track. Growth follows demonstrated impact, not tenure alone. A typical engineering ladder has clear levels with defined expectations for scope, quality, and mentorship. Moving up usually requires owning outcomes end to end rather than completing assigned tickets.
About the company
Graphcore has built a new type of processor for machine intelligence to accelerate machine learning and AI applications for a world of intelligent machines