
Site Reliability Engineer Lead, Debug
Job description
About the role
This position sits at the intersection of hardware and software, requiring a unique blend of technical leadership and hands-on investigation across silicon, firmware, software, and infrastructure layers. The successful hire will own the design and execution of observability and reliability practices specifically for the AI hardware development environment. You will establish technical direction for the Debug and Site Reliability Engineering team, driving initiatives that remove systemic friction and enable long-term engineering success. A core part of this role involves mentoring engineers and shaping incident response processes to ensure the infrastructure supporting AI development remains robust and performant. Systems thinking is the fundamental skill required to navigate the complex interactions between RISC-V architectures, AI accelerators, and distributed compute clusters. Understanding and improving the on-call and incident response lifecycle is integral to maintaining high availability for critical development workflows. Ultimately, this role defines the operational health and reliability standards for the entire engineering infrastructure.
Key facts
What you'll do
Establish reliability, observability, and operational health for engineering infrastructure that supports AI hardware and software development through site reliability practices.
Mentor engineers within the Debug & Site Reliability Engineering team while establishing technical direction for key initiatives that enable long-term operational success.
Drive the implementation of scalable engineering tools and automation using languages such as Python, C++, Go, or Bash to remove friction from development workflows.
Utilize observability platforms like Prometheus, Grafana, OpenTelemetry, and ELK to monitor large-scale production systems and derive actionable insights.
Design and maintain the infrastructure and CI/CD pipelines that underpin the AI compute clusters essential for hardware and software validation.
Develop and refine automated debugging workflows to handle the complex interactions between silicon, firmware, and software in RISC-V and AI accelerator environments.
Collaborate extensively with cross-functional teams to ensure system reliability improves through automation, operational excellence, and shared best practices.
Conduct operational reviews and analyze past incidents to implement structured post-incident thinking and prevent future occurrences.
Build and maintain deep systems knowledge across networking, storage, distributed computing, and hardware/firmware interfaces to solve intricate reliability challenges.
Champion the creation of custom observability and monitoring strategies tailored to the specific needs of AI accelerator debug and validation workflows.
Ensure the infrastructure supports rapid development cycles while maintaining the stability and performance required for cutting-edge AI research.
Lead on-call rotations and incident response processes, ensuring calm and effective handling of critical production issues.
Requirements
The posting states a bachelor's degree requirement; you must You must possess experience spanning a decade or more in building and operating complex software, infrastructure, site reliability, or systems engineering environments.
Your expertise must include extensive debugging capabilities across operating systems, networking, distributed services, hardware, and firmware interactions.
You must demonstrate the ability to develop scalable engineering tools and automation using Python, C++, Go, Bash, or similar programming languages.
You must have hands-on experience with observability platforms such as Prometheus, Grafana, OpenTelemetry, and ELK stacks to monitor large-scale production systems.
You must have a proven track record of driving reliability improvements through automation, operational excellence, and cross-functional collaboration.
You must be comfortable working with export-controlled technology and understand that eligibility for U.S. export access may be a factor in employment.
You must meet the degree requirement as explicitly stated in the official job listing.
Nice to have
Only items explicitly stated as preferred in the source material are considered nice to have; no additional assumptions are made.
Practical notes
The offer may depend on eligibility for U.S. export-controlled technology access.
A degree is required as stated in the listing.
Typical interview steps
Platform interviews usually include an infrastructure scenario, a scripting or coding exercise, and operational questions. Candidates may be asked to design a deployment pipeline or debug an outage. Incident experience and an automation mindset are tested. Interviewers often ask about a past outage and how you handled it. Structured post-incident thinking, not heroics, is what they look for.
Good to know
Debugging complex hardware and software interactions requires deep systems knowledge across multiple layers, including RISC-V CPU designs and AI accelerators.
Large-scale distributed infrastructure for AI compute clusters depends on automated debugging workflows and cross-functional collaboration.
Custom observability and monitoring strategies are essential for RISC-V CPU designs and AI accelerators.
Questions to ask
Useful questions for the interview include inquiring about what a typical week looks like, how work is assigned, what tools the team uses, and how feedback is processed.
It is also reasonable to ask how the role has changed recently and what the team wishes it had known when joining.
Questions regarding the manager's priorities are especially valued during the interview process.
Career growth
Platform careers grow from engineer to senior, staff, and platform lead roles, with some individuals moving into SRE leadership or cloud architecture positions. Breadth across networking, storage, and reliability becomes increasingly important at senior levels. Platform careers reward breadth and calm under pressure, where experience automating your own work serves as the strongest signal for advancement to senior roles.