Site Reliability Engineer
Job description
Site Reliability Engineer at D Matrix.
About the role
You will own the full lifecycle of critical infrastructure systems that power generative AI workloads, spanning on-premises GPU clusters, co-location facilities, and public cloud environments. Your daily work will involve hands-on operations where you build, configure, and maintain the platform services that every engineering team depends on. You will partner directly with senior DevOps leadership to ensure CI/CD pipelines run reliably and efficiently. This role grants high ownership over systems that directly enable silicon validation, AI research experiments, and customer hardware integrations. You will design and implement automation to eliminate manual toil while improving system resilience. The position requires rapid troubleshooting during incidents and clear communication during both normal operations and emergency response. This is a 6-month contract with potential conversion to full-time, offering a direct impact on the company's infrastructure foundation.
Key facts
What you'll do
Infrastructure Operations
- Own reliability and availability of assigned infrastructure domains: colo server fleets, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and customer-facing platform services.
- Perform hands-on infrastructure work: server provisioning, OS configuration, network setup, storage management, and hardware troubleshooting from bare metal up.
- Support and operate high-speed interconnect environments - InfiniBand, RoCE, or high-speed Ethernet - in lab and colo settings.
- Conduct capacity planning and hardware lifecycle management for assigned infrastructure domains.
Automation & Infrastructure as Code
- Own IaC and configuration management (Terraform, Ansible) for your infrastructure domains - all provisioning and changes through code, not manual steps.
- Build automation to eliminate toil: host lifecycle management, fleet health checks, auto-remediation workflows, and self-service tooling for engineering teams.
- Develop networking automations for cluster interconnects, VLAN management, and lab network configurations.
- Contribute to shared IaC modules and automation libraries used across the SRE team.
Observability & Incident Response
- Design and maintain monitoring dashboards, alerting, and SLIs (Prometheus/Grafana, DataDog) for your infrastructure domains - ensuring signal quality and actionable alerts.
- Participate in on-call rotation; triage and resolve incidents from bare metal to application layer, distinguishing infrastructure faults from software or hardware product issues.
- Produce high-quality RCA reports for P0/P1 incidents with root cause analysis and tracked action items.
- Detect performance issues, recommend solutions, and implement fixes that permanently improve system reliability.
Customer & Platform Services
- Support and operate platform services used by external customers for hardware and software deployment collaboration with d-Matrix.
- Ensure QoS and uptime commitments for customer-facing environments; escalate reliability risks proactively.
- Document platform configurations, access procedures, and operational runbooks for customer environments.
Documentation & Collaboration
- Maintain high-quality runbooks, architecture diagrams, and troubleshooting guides - documentation is part of the job, not an afterthought.
- Partner with the DevOps team to ensure infrastructure reliability supports CI/CD pipeline performance and developer experience.
- Serve as a technical resource for engineering teams - sharing operational knowledge and raising infrastructure risks early.
Requirements
- Bachelor's or Master's in Computer Science, Electrical Engineering, or a related field (or equivalent experience); 5+ years in SRE, infrastructure engineering, or systems administration.
- Strong Linux systems knowledge: networking, storage, systemd, package management, kernel parameters, and performance debugging.
- Experience with cloud infrastructure on AWS, Azure, and GCP; deep understanding of at least one of these platforms.
- Hands-on expertise with container orchestration, specifically Kubernetes, including cluster installation, upgrades, and operations.
- Fluency in scripting and automation using Python or Bash to manage infrastructure at scale.
- Demonstrated ability to operate and maintain high-speed networking in data center environments including TCP/IP, routing, switching, and firewall concepts.
- Experience managing distributed systems and troubleshooting performance bottlenecks in complex environments.
- Commitment to on-call rotation and incident response with a mindset focused on service reliability and customer impact.
Practical notes
This is a 6-month contract with potential conversion to full-time. It is a hands-on, high-ownership role. You will build, operate, and automate real infrastructure - not manage tickets. The systems you keep running directly enable silicon development, AI/ML research, and customer success. You will work closely with the senior DevOps lead team, whose CI/CD pipelines and automation layer depend on the infrastructure you operate. You will also support customer-facing environments where d-Matrix partners collaborate on hardware and software deployments.