Staff Engineering Operations Technical Program Manager
Job description
About the role
This role is centered on delivering advanced operational and diagnostic support for Graphcore's Arm-based hardware platforms within lab and data center environments. You will be responsible for leading hardware bring-up, validation, and troubleshooting activities for complex AI compute platforms. The position requires a focus on ensuring the reliability, stability, and high performance of next-generation AI hardware systems. Collaboration with various engineering teams - including hardware, systems, and data center operations - is essential to identify issues, implement solutions, and improve hardware design and deployment processes. You will play a key role in maintaining the operational excellence of Graphcore's hardware infrastructure, supporting engineering development cycles, and ensuring smooth integration of new hardware platforms.
Key facts
What you'll do
- Lead advanced troubleshooting efforts for server blades, motherboards, power systems, and rack-scale infrastructure to resolve hardware failures efficiently.
- Support the hardware bring-up process, including validating new components, testing firmware interactions, and ensuring hardware compatibility.
- Diagnose system failures related to thermal behavior, power issues, network configuration, BIOS/BMC errors, and other hardware-related problems.
- Collaborate closely with server engineering teams to perform root cause analysis, identify design flaws, and recommend improvements to hardware architecture.
- Assist in the deployment of new hardware platforms by conducting structured validation, qualification, and testing procedures to ensure readiness for production.
- Interface with facilities and data center teams to understand environmental factors such as cooling, power distribution, and physical infrastructure that impact hardware reliability.
- Develop, document, and maintain standard operating procedures, troubleshooting guides, and validation checklists to support ongoing hardware diagnostics and deployment activities.
- Mentor junior technicians and engineers on hardware troubleshooting techniques, diagnostic procedures, and best practices for hardware diagnostics.
- Participate in on-call rotations or provide off-hours support during critical engineering milestones, hardware releases, or system failures to ensure minimal downtime and rapid resolution.
- Contribute to continuous improvement initiatives by analyzing failure data, identifying trends, and recommending process or design changes to enhance hardware robustness.
- Support cross-functional projects aimed at hardware upgrades, system scaling, and infrastructure improvements to meet evolving engineering requirements.
Requirements
- Bachelor's degree in Electrical Engineering, Computer Engineering, Computer Science, or a related technical field.
- 12 years of hands-on experience with server hardware architectures, board-level debugging, and hardware diagnostics.
- Proven ability to analyze system logs, hardware telemetry data, and power/thermal metrics to identify root causes of hardware failures.
- Extensive experience working with high-performance computing (HPC) systems, AI compute platforms, or rack-scale infrastructure.
- Strong understanding of server hardware components, firmware interactions, and hardware validation processes.
- Excellent collaboration skills to work effectively within fast-paced, multidisciplinary engineering teams.
- Strong written and verbal communication skills to document procedures, create reports, and communicate findings clearly.
- Ability to work independently, prioritize tasks, and manage multiple troubleshooting projects simultaneously.
- Familiarity with hardware diagnostic tools, oscilloscopes, logic analyzers, and other debugging equipment is preferred.
- Knowledge of data center environmental factors, including cooling systems, power distribution, and physical infrastructure, is advantageous.
Nice to have
- Experience supporting prototype or pre-production hardware bring-up, including early-stage validation and testing.
- Familiarity with data center infrastructure, including liquid cooling systems, power distribution units, and environmental monitoring.
- Experience using scripting languages such as Python or Bash to automate hardware validation, diagnostics, or data collection tasks.
- Exposure to structured failure analysis methodologies, reliability engineering practices, and design for testability principles.
- Knowledge of firmware development, BIOS, BMC, or embedded system firmware is a plus.
- Understanding of industry standards related to hardware reliability and safety.
Skills & tools
- Arm-based hardware platforms
- Server hardware architectures
- Board-level debugging techniques and tools
- System logs, hardware telemetry, power and thermal metrics analysis
- HPC systems, AI compute platforms, rack-scale infrastructure
- Automation scripting with Python, Bash (preferred but not mandatory)
- Diagnostic equipment such as oscilloscopes, logic analyzers, and hardware debugging tools
Practical notes
Graphcore offers a comprehensive benefits package that includes medical, dental, and vision coverage, along with flexible spending accounts (FSAs) and health savings accounts (HSAs). The company provides disability and life insurance, a 401(k) retirement plan, commuter benefits, wellness programs, and an Employee Assistance Program (EAP). The organization supports flexible working arrangements and is committed to fostering an inclusive and diverse work environment. Reasonable accommodations for interviews and onboarding are available upon request to ensure accessibility for all candidates. The role is based on-site in Austin, Texas, and requires physical presence at the office or data center facilities. The company values technical excellence, collaboration, and continuous improvement, making this an excellent opportunity for experienced hardware troubleshooting professionals seeking to impact AI hardware development.