Lab Support Engineer
Job description
About the role
You will maintain the operational health of our hardware labs and silicon development environments. This role is centered on providing robust technical assistance to both internal engineering teams and external partners who rely on our critical infrastructure. A core responsibility is to ensure that all systems remain secure, functional, and available for high-priority development work. You will act as a first line of defense against operational issues, using your skills to resolve incidents efficiently. The position requires a proactive mindset to identify potential problems before they escalate into major disruptions. You will work closely with hardware specialists to support the lifecycle of complex server equipment. This role is integral to the success of the Hardware Labs and Silicon Development teams.
Key facts
What you'll do
- Triage and resolve L1 and L2 support tickets using direct communication and our internal tracking system to restore service quickly.
- Oversee the daily operation of our server fleet, ensuring uptime and assisting hardware lab teams with routine administrative tasks.
- Execute physical server installations, perform hardware maintenance, and run diagnostics to verify the integrity of components.
- Create and update documentation for technical resolutions to maintain a current and useful internal knowledge base.
- Administer Linux-based systems, applying configurations and updates to support ongoing development projects effectively.
- Monitor hardware health indicators and assist in diagnosing issues related to server racks and peripheral devices.
- Implement standard security procedures to safeguard infrastructure and respond to potential incidents in a controlled manner.
- Collaborate with network engineers to manage connectivity, addressing issues related to routing, subnets, and VLANs.
- Support the deployment of services by configuring reverse proxies, load balancers, and web servers as required by the team.
- Utilize infrastructure-as-code tools to automate server provisioning and maintain consistent environments across the lab.
Requirements
- Demonstrate proficiency in Linux administration, with a specific focus on Debian and RedHat distributions and their respective ecosystems.
- Possess experience with core networking concepts, including routing, subnets, VLANs, VPNs, and wireless connectivity principles.
- Show knowledge of server and desktop hardware, including rack mounting techniques, PDUs, BIOS/firmware updates, and BMC/Out-of-Band management protocols.
- Exhibit the ability to manage infrastructure using automation tools such as Ansible or Puppet to ensure reliability and scalability.
- Apply strong problem-solving skills to complex technical issues while maintaining a customer service-oriented approach.
- Hold the capability to work independently and as part of a collaborative team in a fast-paced engineering environment.
- Adhere to strict security and compliance standards to protect company data and intellectual property.
- Communicate effectively with both technical and non-technical stakeholders to clarify requirements and status updates.
Nice to have
- Experience diagnosing performance bottlenecks related to CPU, RAM, storage, or network throughput to optimize system behavior.
- Familiarity with monitoring stacks such as Prometheus, Zabbix, Grafana Mimir, or Open Telemetry for observability and alerting.
- Ability to configure reverse proxies, load balancers, and web servers like Nginx or HAProxy to manage traffic routing.
- Practical knowledge of container orchestration and frameworks like Kubernetes, Docker, or containerd for managing application deployment.
- Python scripting skills for API interaction, data processing, or utility development to automate repetitive tasks.
Practical notes
Graphcore is part of the SoftBank Group and maintains an inclusive hiring process. We offer flexible interview arrangements and encourage candidates to request any necessary reasonable adjustments during the application process. The role is based in Bristol, United Kingdom, and is offered as a full-time engagement. This position is part of the Hardware Labs and Silicon Development team, working directly on the infrastructure that powers advanced silicon development. Success in this role requires a high level of responsibility and attention to detail. The work environment is technical and demanding, requiring dedication and a passion for infrastructure reliability.