Staff Systems Engineer
Job description
About the role
Graphcore is a SoftBank Group company focused on building advanced AI compute hardware and infrastructure. This role provides technical expertise to maintain, validate, and troubleshoot Arm-based systems within lab and data center settings. The Staff Systems Engineer owns the deep technical ownership of complex server infrastructure, driving the reliability and performance of critical hardware platforms. You will serve as the definitive expert responsible for diagnosing the most intricate hardware issues across the entire lifecycle of deployment. This position requires a proactive approach to identifying systemic risks before they escalate into critical failures. You will partner directly with cross-functional teams to ensure hardware stability aligns with ambitious product goals. The work involves a continuous cycle of investigation, validation, and knowledge transfer to elevate the overall robustness of the platform.
Key facts
What you'll do
- Execute advanced diagnostics and perform complex repairs on server blades, motherboards, power units, and rack-scale systems to restore operational integrity.
- Drive hardware bring-up activities by assisting in testing firmware and validating components against strict quality and performance benchmarks.
- Investigate and resolve system failures stemming from thermal anomalies, power irregularities, networking faults, BIOS misconfigurations, and BMC operational issues.
- Lead root cause analysis sessions in collaboration with server engineering teams to transform findings into actionable design improvements.
- Oversee the validation and qualification phases for new hardware deployments to ensure they meet stringent reliability standards.
- Liaise with facilities teams to monitor environmental impacts on system performance and recommend corrective actions for optimal conditions.
- Author comprehensive technical documentation, detailed troubleshooting guides, and robust standard operating procedures for repeatable processes.
- Develop the capabilities of junior staff through mentorship on diagnostic techniques and hardware repair workflows to build team resilience.
- Participate actively in on-call rotations to provide immediate support during critical project milestones and urgent hardware launches.
- Analyze telemetry data streams to identify trends and anomalies that indicate potential future hardware degradation or failure modes.
- Implement automated validation scripts to accelerate testing procedures and reduce manual intervention in repetitive tasks.
- Collaborate with vendors and partners to reproduce and resolve issues that require specialized component-level knowledge.
- Ensure all troubleshooting activities adhere to strict safety protocols and data center operational standards.
- Contribute to the architecture of monitoring systems to improve visibility into hardware health and performance metrics.
Requirements
- Hold a Bachelor degree in Electrical Engineering, Computer Engineering, Computer Science, or a related field from an accredited institution.
- Bring a minimum of 10 years of professional experience specifically in board-level debugging and server hardware architecture.
- Demonstrate a proven ability to isolate hardware faults using system logs, telemetry data, and power or thermal metrics effectively.
- Show practical background working with High-Performance Computing (HPC) systems, AI compute platforms, or rack-scale infrastructure.
- Exhibit strong communication skills to articulate technical findings clearly to both technical and non-technical stakeholders.
- Function effectively within collaborative engineering groups that operate under tight deadlines and high-stakes environments.
- Possess the legal right to work in the United States without sponsorship requirements for this position.
- Maintain the ability to pass background checks and meet the necessary compliance standards for employment in a Milpitas facility.
- Be willing to work from the Milpitas office location consistently to ensure seamless integration with the team.
Nice to have
- Possess background in bringing up prototype or pre-production hardware to refine designs before mass deployment.
- Show knowledge of data center infrastructure, such as power distribution and liquid cooling systems, to understand holistic environments.
- Demonstrate proficiency in Python, Bash, or other automation scripts for validation tasks to improve efficiency.
- Have familiarity with reliability engineering and formal failure analysis processes to enhance systematic problem-solving.
Practical notes
This role operates on a Full-time basis with standard employment terms applicable to the location. The position is based in Milpitas, California, United States, requiring consistent presence at the office. Graphcore provides comprehensive benefits including medical, dental, and vision insurance, 401(k) retirement plans, disability and life insurance, FSAs, HSAs, commuter benefits, and an Employee Assistance Programme. The company supports flexible working arrangements and provides reasonable adjustments for the interview process. Graphcore is an equal opportunity employer.