Senior/Staff Site Reliability Engineer
Job description
About the role
You will advance the state of our operations by implementing SRE best practices with a strong focus on users, monitoring, and automation across our infrastructure. You will design, build, and operate our on-premises data center environment to directly support the rapid growth of our Machine Learning team. You will build and maintain highly-secure on-premises environments that adhere to NIST and ISO standards, ensuring robust compliance and operational integrity. You will integrate our on-premises datacenter environments with existing cloud infrastructure to create a seamless and efficient hybrid cloud experience. You will improve the reliability and resilience of our infrastructure by driving root-cause analysis and reviewing gaps in designs and implementations. You will participate in platform on-call rotations and assist with urgent incident response as a key member of the operations team. You will own the end-to-end reliability of critical production systems, ensuring they meet stringent uptime and performance targets for our pathology platform.
Key facts
What you'll do
Advancing the state of our operations by implementing SRE best practices - focusing on users, monitoring, and automation.
Designing, building, and operating our data center to support our rapidly growing Machine Learning team.
Building highly-secure on-premises environments handling NIST/ISO standards.
Integrating on-premises datacenter environments with existing cloud infrastructure to create a seamless hybrid cloud environment.
Improving the reliability and resilience of our infrastructure through root-cause analysis and reviewing gaps in designs, and implementations of our infrastructure.
Participating in platform on-call rotations and assisting with urgent incident response.
Eliminating toil by automating infrastructure operations through scripting and configuration management tools such as Ansible and RedFish.
Building monitoring infrastructure with modern observability tools including Datadog, Grafana, and Prometheus.
Managing critical production infrastructure and demonstrating experience with incident response, scaling, and rapid growth related challenges.
Maintaining physical hardware stacks in production settings using iDRAC, IPMI, Nvidia UFM, and Juniper Systems.
Some experience and opinions on virtualization, containerization, or container orchestration platforms such as EKS-Anywhere, ClusterAPI, and KVM.
You are opinionated on storage solutions and how they can be optimized for high performance workloads including Quobyte, S3, FSx, and EFS.
Occasional travel to onsite Datacenter location(s) as required for operational duties.
A bachelor's degree in Computer Science or equivalent experience is required for this role.
Requirements
You must possess 8+ years of relevant experience in Site Reliability Engineering or related infrastructure roles.
You have familiarity with modern datacenter network designs and comfort operating across network layers including physical, L2/L3, and overlay networks.
You have administered physical hardware stacks in production settings using iDRAC, IPMI, Nvidia UFM, and Juniper Systems.
You have some experience and opinions on virtualization, containerization, or container orchestration platforms such as EKS-Anywhere, ClusterAPI, and KVM.
You are opinionated on storage solutions and how they can be optimized for high performance workloads such as Quobyte, S3, FSx, and EFS.
You automate everything to eliminate toil, leveraging scripting and configuration management tools like Ansible and RedFish.
You have built monitoring infrastructure using modern observability tools including Datadog, Grafana, and Prometheus.
You have operations experience managing critical production infrastructure, including incident response, scaling, and rapid growth challenges.
You hold a bachelor's degree in Computer Science or possess equivalent practical experience.
You have an insatiable intellectual curiosity and the ability to learn quickly in a complex technical space.
You are comfortable making decisions in ambiguous situations while owning end-to-end outcomes for critical systems.
You communicate clearly and collaborate effectively with cross-functional teams in a fast-paced environment.
You are willing to perform occasional travel to onsite datacenter locations as needed for operational requirements.
You are eligible to work in the United States and require no sponsorship for this role at this time.
Practical notes
Hours: Full-time (40 hours per week).
Travel: Occasional travel to onsite Datacenter location(s).
Visa: Eligible to work in the United States and require no sponsorship for this role at this time.
Deadline: No application deadline specified.