Skip to main content
etched careers
etched logo

Head of Platform Product Reliability

etchedUSAFull Time2mo ago
AIHTMLOperationsSupportTalentGrowthStrategyEngineeringInfrastructurePlatformReliabilityTesting

Continue with your CV

PDF, Word, or image. Under 10MB.

Job description

About Etched

Etched is building hardware for frontier intelligence. We co-design chips, racks, software, and manufacturing to deliver best-in-class throughput and latency across both prefill and decode workloads. Our first products are heavily focused on inference. Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history.

Job Summary

We are seeking a highly technical and execution-focused Head of Platform Product Reliability to lead reliability engineering across Etched's server, rack, and datacenter platform products.

This role owns system-level product reliability from architecture through fleet deployment. You will define reliability strategy, qualification methodologies, accelerated stress testing programs, failure analysis processes, and long-term reliability standards for complex AI infrastructure systems. This team focuses specifically on product reliability engineering for platform hardware and deployed systems - ensuring every Etched product ships with the reliability profile that enterprise and hyperscale customers demand.

You will work cross-functionally with Platform Engineering, Mechanical Engineering, Thermal, Firmware, Manufacturing, Supply Chain, Datacenter Operations, and Program teams to ensure Etched products achieve exceptional reliability at scale.

Key Responsibilities

- Define and own the end-to-end reliability strategy for AI servers, accelerator platforms, rack systems, and datacenter infrastructure, from design requirements through field deployment

- Establish reliability requirements, qualification standards, and validation methodologies that scale across product generations

- Build and institutionalize reliability engineering processes spanning the full product lifecycle:

- EVT / DVT / PVT qualification gates and exit criteria

- Accelerated life testing (ALT) and accelerated stress testing (AST)

- Environmental testing: temperature, humidity, altitude, contamination

- HALT / HASS programs for design margin and production screening

- Vibration, shock, and transportation stress testing

- Power cycling, thermal cycling, and long-duration soak testing

- Lead root-cause investigations for reliability failures surfaced during development, manufacturing, and field deployment, driving corrective actions across hardware, firmware, thermal, and mechanical domains

- Develop comprehensive system reliability models including MTBF projections, FIT rate analysis, Weibull lifetime modeling, component derating methodologies, and reliability growth tracking

- Ensure reliability is considered early, partnering with Platform Engineering architects and design leads so reliability requirements shape decisions before they become expensive to change

- Work closely with ODMs, JDMs, contract manufacturers, and component suppliers to validate and enforce long-term platform reliability commitments

- Build fleet reliability infrastructure: telemetry analysis pipelines, field feedback loops, and monitoring frameworks that give Etched visibility into deployed system health at scale

- Drive reliability signoff criteria and lead product release readiness reviews across engineering and program teams

- Build and lead a high-performing product reliability engineering organization - hiring, developing, and retaining technical talent as the company scales

You may be a good fit if you have (Must-have qualifications)

- BS, MS, or PhD in Electrical Engineering, Mechanical Engineering, Reliability Engineering, or a related technical field

- 10+ years of reliability engineering experience in hardware-centric organizations, with meaningful time spent on complex systems rather than component-level work

- Experience leading reliability programs for one or more of:

- AI accelerator or GPU-class compute systems

- Hyperscale or cloud server infrastructure

- Networking platforms, storage systems, or rack-scale infrastructure

- Deep understanding of system-level failure mechanisms - including thermal, power delivery, mechanical, and connector/interconnect failure modes - and how design decisions affect long-term field reliability

- Hands-on experience with FMEA, Weibull analysis, HALT/HASS, qualification planning, failure analysis methodologies, and reliability statistics and modeling

- A track record of driving cross-functional root-cause investigations in fast-moving hardware organizations where schedule pressure is real and accountability is high

- Strong technical judgment - capable of making defensible tradeoffs between reliability targets, cost, schedule, and performance without losing sight of customer expectations

- Excellent communication skills and the credibility to influence design decisions with engineering leads, program managers, and executive stakeholders

Strong candidates may also