Principal Hardware Diagnostics Engineer
Job description
About the role
Graphcore is hiring a Principal Hardware Diagnostics Engineer to build diagnostics software that monitors hardware health and troubleshoots system-level problems across AI infrastructure platforms. You will create diagnostics agents, tools, and analytics frameworks that help engineers and automated systems detect, isolate, and fix hardware faults in blade servers and rack-scale clusters. In this role, you will own the design and implementation of diagnostic capabilities that span from firmware interfaces to production observability. You will act as the bridge between low-level hardware behavior and high-level system reliability, ensuring that faults are caught early and resolved quickly. The position requires deep collaboration with cross-functional engineering teams to embed diagnostics directly into the control plane and validation flows. You will drive the creation of tools that turn raw telemetry into actionable insights for both human experts and automated remediation systems. This role is critical for maintaining the robustness and predictability of Graphcore's AI infrastructure platforms at scale.
Key facts
What you'll do
- Build automated hardware diagnostics solutions for blade-level servers and rack-scale AI systems using scalable software architectures.
- Architect diagnostic agents, monitoring tools, and analytics frameworks to capture hardware telemetry across system boundaries and lifecycle stages.
- Partner with hardware teams to embed low-level diagnostic modules into monitoring systems, ensuring coverage from boot to production.
- Create tools that identify hardware health conditions and pinpoint failures with high precision to reduce mean time to repair.
- Develop diagnostic modules for internal validation and production data center operations, supporting both pre-silicon and post-deployment scenarios.
- Supply detailed hardware fault data to system engineers to speed up troubleshooting and improve long-term product reliability.
- Establish remediation workflows and insights for hardware fault scenarios across nodes and clusters, including escalation paths and runbooks.
- Work with firmware, networking, and cloud platform teams to integrate diagnostics throughout the system stack, from drivers to orchestration layers.
- Define and implement tests that validate diagnostic accuracy, performance, and robustness under real-world failure modes.
- Contribute to open standards and internal frameworks that standardize how diagnostics are modeled, collected, and consumed across Graphcore platforms.
Requirements
- Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, or a related field with a strong focus on systems and hardware interaction.
- Solid software engineering background in Python, C++, or C# with a proven track record of writing clean, maintainable, and performant code.
- Hands-on experience building diagnostics or monitoring systems for hardware platforms, including instrumentation, data collection, and failure analysis.
- Background working with distributed systems or cloud infrastructure, including an understanding of scalability, fault tolerance, and observability.
- Strong familiarity with Linux environments and system-level diagnostics tools, including logs, metrics, traces, and kernel interfaces.
- Experience partnering with CM/ODM partners on manufacturing diagnostics and fault isolation, including test development and yield analysis.
- Ability to work independently and collaboratively in a fast-paced, matrixed engineering environment with ambiguous problems.
- Strong written and verbal communication skills to translate complex hardware issues into clear diagnostics and actionable reports.
Skills & tools
- Python, C++, C#
- Linux system-level diagnostics tools
- Distributed systems and cloud infrastructure
- Hardware telemetry and monitoring frameworks
Nice to have
- Experience with AI or machine learning infrastructure platforms and their failure modes.
- Familiarity with Graphcore IPU hardware or other accelerator architectures.
- Background in firmware or low-level system bring-up and diagnostics.
- Knowledge of observability standards and tracing frameworks used in large-scale deployments.
- Experience with automated test frameworks and continuous validation pipelines.
Practical notes
- Location options include Austin, Texas, United States or Milpitas, California, United States, and may influence team assignment and office-specific policies.
- This is a full-time engagement with opportunities for career growth and technical impact within the Systems Engineering and Platform Validation organization.
- Candidates must be authorized to work in the United States without sponsorship for this role.
- All employment decisions are subject to Graphcore's standard policies, background checks, and eligibility verification processes.