Software Infrastructure Kubernetes Engineer
Job description
About the role
You will own the reliability and performance of our Kubernetes-based infrastructure pipelines. You will design and iterate on deployment workflows so teams can move safely and with confidence. Your work directly shapes how AI experiments are run, observed, and iterated upon by researchers. You will join the Software Infrastructure team as a key contributor to their platform and tooling. In this role, you will develop essential tools and services that empower our broader software engineering organization. Your contributions will enhance the build, test, deployment, and productisation processes of our Machine Learning Software components. You will work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed systems and large-scale compute. This position operates at the intersection of infrastructure engineering and AI workloads, ensuring the platform meets the demands of cutting-edge research.
Key facts
What you'll do
Design intake pipelines that capture requirements and constraints specific to AI workloads and cluster usage.
Build and maintain container images and Kubernetes manifests that adhere to production-grade standards and security policies.
Review changes to infrastructure code with a strong emphasis on correctness, clarity, and long-term maintainability.
Ship updates to cluster environments using controlled delivery strategies and implement robust monitoring for regressions.
Partner with algorithm and systems teams to enable scalable and efficient machine learning workflows across the organization.
Define and operationalize patterns that reduce manual intervention and improve the overall reliability of the service platform.
Automate testing and validation for infrastructure changes across development, staging, and production environments.
Guard cluster resources and tune scheduling policies to support demanding AI training and inference tasks effectively.
Deploy and maintain the Kubernetes infrastructure that underpins the development, testing, and scaling of Graphcore hardware and its software stack.
Develop, own, and maintain the tools and services that support the software organization in their daily workflows.
Implement observability practices to ensure performance metrics and logs are actionable for rapid incident investigation.
Collaborate with cross-functional teams to translate operational requirements into infrastructure improvements.
Ensure that all deployed configurations follow best practices for version control, auditability, and disaster recovery.
Contribute to the evolution of the platform by proposing and prototyping improvements to the CI/CD and deployment pipelines.
Requirements
You have hands-on experience with Kubernetes and related orchestration tooling in production environments.
You understand container image builds, registries, and how they integrate with deployment pipelines and runtime environments.
You know YAML-based configuration and Helm-like templates for defining services, roles, and networking resources.
You are comfortable working with command-line tools and editing configuration files over SSH in remote server environments.
You have used monitoring or logging systems such as Prometheus and Grafana to investigate incidents in production.
You have at least three years of experience managing infrastructure as code repositories using tools like Git.
You are familiar with the principles of high-availability clusters and disaster recovery strategies for critical services.
You have a strong understanding of Linux operating systems, networking concepts, and process management in distributed systems.
Nice to have
Experience with high-performance computing or AI-related software stacks is a plus for this role.
Practical notes
This is a full-time role based in London, UK.
The compensation for this position is £70,000 per year for a candidate with 3 years of experience.
Please apply only through the channels specified by Graphcore recruitment teams.
This role involves working with cutting-edge AI hardware and software platforms.
The position may require interaction with HPC workloads and distributed computing environments.
Relevant security clearance or compliance checks may be required depending on project scope.
Candidates must be legally authorized to work in the United Kingdom.
Travel requirements are not applicable for this role.
Visa sponsorship may be considered for eligible candidates under UK immigration rules.
Deadlines for application review are managed by the recruitment team on a rolling basis.