Staff System Tests & Diagnostics Engineer
Job description
System Tests & Diagnostics Engineer at Graphcore.
About the role
This role develops and extends diagnostics and stress tools for next-generation AI SoCs. Collaboration with Arm engineers shapes Graphcore-specific diagnostic capabilities.
Software engineers turn product ideas into working code. Engineers work in small teams, review each other's work, and ship in small batches. Most teams follow agile practices such as sprints and daily standups. Engineers also write tests, fix bugs, and improve performance. The field values clear communication as much as technical skill. Engineers spend part of every week on planning, code review, and debugging, not just writing new code. The ability to explain a technical decision in plain words separates strong engineers from the rest.
Key facts
What you'll do
Exposure of intermittent hardware failures is handled by system-level diagnostics. Automated execution across server and rack-scale validation environments is supported.
Scalable diagnostic applications are designed with configuration-driven approaches. Support for multiple silicon revisions, platforms, and execution environments is enabled through parameterization and workload variation.
Analysis of diagnostic output drives improvements in fault isolation methodologies.
Reproduction and investigation of difficult hardware failures are enabled for engineers through developed software.
Requirements
A degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical discipline is required.
Candidates must bring 10+ years of experience developing hardware diagnostics, system validation software, firmware validation tools, or low-level systems software.
Strong software development skills in Python and C/C++ are necessary.
Robust Linux systems experience is required.
Understanding of modern server platform architecture, including CPU, memory, PCIe, storage, networking, firmware, operating system, and device drivers, must be demonstrated.
Experience debugging hardware/software interactions across multiple platform layers is essential.
Background in developing diagnostics, stress testing, or hardware validation software is required.
Experience with server platforms, embedded systems, or SoC validation is necessary.
Strong analytical and debugging skills are mandatory.
Collaboration across hardware, firmware, software, and validation teams must be effective.
Excellent communication and problem-solving skills are required.
Practical notes
This role operates in a US-based environment with standard employment eligibility requirements. Typical interview steps
Hiring for engineering roles usually starts with a recruiter screen, followed by one or two technical rounds. Candidates often solve a coding problem, discuss past projects, and answer system design questions. Some loops include a take-home task. Final rounds typically cover team fit and give candidates a chance to ask questions. Interviewers look for how you break down an unfamiliar problem, not just whether you reach the answer. Practicing a few problems aloud and reviewing your own past projects are the best preparation.
Good to know
Diagnostics engineers build tools that validate hardware at scale. These roles work across firmware, software, and hardware layers. Silent data corruption testing is a common focus in advanced compute systems. Server platform architecture defines many validation constraints. Configuration-driven approaches enable scalable testing. Hardware telemetry feeds diagnostics workflows. Debugging spans multiple abstraction layers.
Questions to ask
Worth asking in any interview: how the team measures success, who the role works with daily, what the onboarding looks like, and what the company is trying to achieve this year. Asking what past hires did well is a strong final question. Keep the list short and pick the questions that matter most to you.
Career growth
Engineering careers usually progress from individual contributor to senior, staff, and principal levels. Some engineers move into management and lead teams of five to twenty people. Others stay on the technical track. Growth follows demonstrated impact, not tenure alone. A typical engineering ladder has clear levels with defined expectations for scope, quality, and mentorship. Moving up usually requires owning outcomes end to end rather than completing assigned tickets.