Linux Engineering Lead
Job description
About the role
Graphcore is actively searching for a Linux Engineering Lead to define and manage the Linux platforms that empower its engineering teams across the company. This role is a technical leadership position that requires equal parts strategic thinking and hands-on involvement, where you will remain deeply engaged in troubleshooting, automation, and day-to-day system support. You will be entrusted with owning the incident response processes and driving operational improvements that enhance reliability and efficiency. In this capacity, you will design and enforce the standards and practices that govern how Linux systems are managed, deployed, and maintained organization-wide. The successful candidate will work closely with cross-functional groups to ensure that the infrastructure keeps pace with ambitious engineering goals. You will champion the adoption of modern practices such as Infrastructure-as-Code and GitOps to bring consistency and speed to operations. Ultimately, this role will shape the long-term vision and direction of Linux platform engineering at Graphcore, ensuring the platform is robust, scalable, and aligned with business needs.
Key facts
What you'll do
- Guide and develop a team of Linux engineers, providing mentorship, career direction, and technical leadership.
- Assume full ownership of reliability, performance, and scalability for engineering Linux environments, monitoring and addressing risks proactively.
- Drive the adoption of Infrastructure-as-Code and GitOps principles across the organization to standardize and simplify operations.
- Build and maintain automation scripts and tools that reduce manual overhead, improve consistency, and accelerate delivery.
- Lead major incident response efforts, coordinating with relevant teams and conducting thorough root cause analysis to prevent recurrence.
- Collaborate closely with software, platform, and infrastructure groups to analyze evolving engineering needs and deliver tailored solutions.
- Define and implement standards, tooling, and processes that allow systems to scale efficiently while maintaining security and compliance.
- Enhance observability, monitoring, and operational visibility across all systems to enable data-driven decision-making.
- Contribute to the strategic roadmap for the Linux platform, helping to prioritize initiatives and allocate resources effectively.
- Serve as a key point of contact for stakeholders seeking reliable and high-performance Linux infrastructure support.
- Mentor junior engineers by providing code reviews, technical guidance, and hands-on assistance when resolving complex issues.
- Evaluate and recommend new tools, technologies, and best practices that can be adopted to improve the efficiency of the platform.
- Partner with SRE and DevOps teams to align on shared goals around availability, resilience, and operational excellence.
- Document processes, configurations, and runbooks to ensure knowledge is shared and systems are easy to manage at scale.
Requirements
- Demonstrate substantial experience managing Linux environments at scale, handling diverse workloads and high-availability scenarios.
- Show solid troubleshooting ability across systems, networking, storage, and applications, with a methodical approach to diagnosing issues.
- Provide a track record of leading engineers, projects, or operational efforts, with evidence of delivering results on time and to a high standard.
- Exhibit strong scripting and automation skills using Python, Bash, or similar languages to solve operational challenges efficiently.
- Bring hands-on experience with Infrastructure-as-Code and configuration management tools such as Ansible, Terraform, or Puppet.
- Show familiarity with Git-based workflows and CI/CD pipelines, understanding how they integrate with infrastructure management.
- Have a background in handling production incidents and driving operational improvements, including post-incident reviews and follow-ups.
- Communicate clearly and manage stakeholders effectively, translating technical details into actionable insights for both technical and non-technical audiences.
Nice to have
- Gain experience supporting AI, HPC, or large-scale engineering environments where performance, scalability, and reliability are critical.
- Familiarity with observability platforms and monitoring systems, including metrics, logs, and tracing solutions.
- Background working alongside platform engineering, SRE, or DevOps teams in highly collaborative settings.
- Knowledge of identity and access management concepts and how they intersect with infrastructure and security.
- Experience building or scaling operational processes that improve efficiency, governance, and auditability.
Practical notes
- Graphcore is part of the SoftBank Group with long-term investment backing.
- Flexible interview arrangements available upon request.
- Inclusive hiring process welcoming candidates from varied backgrounds.