Director, Site Reliability Engineering
Job description
Director, Site Reliability Engineering at D Matrix.
About the role
D Matrix is building a site reliability engineering function from the ground up to support its leadership in AI inference silicon design and manufacturing. The role owner will establish the SRE discipline and define operational excellence across all infrastructure domains. You will own the reliability, observability, and capacity planning for development, validation, and customer-facing deployments. This position requires direct hands-on leadership in setting technical direction and owning critical incident response. You will partner closely with hardware, software, and DevOps teams to ensure infrastructure meets workload requirements. The role is a 6 month contract with a path to full-time conversion for the right candidate. You will build the processes, tools, and culture that allow D Matrix to scale reliably in a competitive environment.
Key facts
What you'll do
- Own the end-to-end establishment of the SRE function at D Matrix, defining the team charter and operational culture from scratch.
- Hire, develop, and retain a team of 3-5 SRE engineers, fostering a culture of ownership, operational excellence, and continuous improvement.
- Define and execute the SRE technical roadmap, including reliability architecture, automation priorities, capacity planning, and on-call model design.
- Serve as the senior technical escalation lead for critical incidents, driving cross-team triage, root cause analysis, and systemic fixes.
- Translate infrastructure health and operational signals into clear narratives for engineering leadership and executive stakeholders.
- Partner with the director of DevOps engineering to align infrastructure reliability with pipeline and automation delivery as a unified platform.
- Direct a dedicated data center and lab technician team, setting priorities and operational standards for on-premises and colocation facilities.
- Establish SRE processes from a zero baseline, including SLIs, SLOs, error budgets, on-call rotations, and incident management frameworks.
- Own 24x7 reliability across colocation, on-premises lab clusters, cloud environments, and customer-facing platform services.
- Design for failure domains, implement progressive delivery, and enforce strict change control at every tier of the infrastructure stack.
- Own the full observability stack from the ground up, including metrics, traces, and logs using tools such as Prometheus, Grafana, and Datadog.
- Evolve incident and problem management into a data-driven discipline with automated triage workflows and pattern detection.
- Drive FinOps and capacity planning as a unified discipline across cloud, colocation, and on-premises tiers, including spend visibility and TCO modeling.
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related technical field or equivalent practical experience.
- 10+ years of experience in infrastructure, platform, or site reliability engineering roles with increasing scope and impact.
- 5+ years of hands-on experience leading SRE or platform teams in building processes, tools, and operational practices.
- Demonstrated experience establishing SRE functions in early-stage organizations where infrastructure is a competitive differentiator.
- Expert-level proficiency in Linux systems administration, networking, and cloud infrastructure on AWS, Azure, and GCP.
- Strong scripting and programming skills in Python, Go, or similar languages for automation and tooling.
- Deep experience with observability tools and practices, including metrics, tracing, and log aggregation and analysis.
- Proven ability to manage on-call rotations, incident response, and post-incident review processes in 24x7 environments.
Nice to have
- Experience with AI accelerator hardware, inference workloads, or data center environments.
- Familiarity with colocation and direct hardware integration in multi-cloud environments.
- Background in HPC or high-throughput compute workloads in cloud and on-premises environments.
Practical notes
This is a 6 month contract role with a path to full-time conversion. The position is based in Santa Clara and may require travel within the site and between facilities as needed. There may be visa sponsorship considerations for qualified candidates. The role reports to the director level and requires availability to respond to critical incidents as part of an on-call rotation.