Senior Site Reliability Engineer
Job description
About the role
We are actively building a robust and scalable infrastructure platform to support a significant expansion in our engineering capacity, targeting growth of over 90 new team members this year. The Senior Site Reliability Engineer II will be a central figure in ensuring the long-term reliability, performance, and observability of the core systems our organization depends on. This position requires deep ownership of the entire system lifecycle, from initial design and deployment through ongoing operations and optimization in production. You will work at the intersection of infrastructure, automation, and process to remove barriers for a large number of engineering teams. Success in this role will be measured by the stability of the platforms you build and the velocity you enable across the engineering organization. You will be instrumental in shaping how we deliver software reliably and efficiently as we scale.
Key facts
What you'll do
- Oversee organizational reliability and observability strategy, defining and tracking SLAs and SLOs while instrumenting systems for comprehensive monitoring and health management.
- Architect, deploy, and sustain our Kubernetes foundation on EKS, including the management of ECS workloads, Helm chart lifecycle, and the configuration of core AWS services such as RDS, Aurora Postgres, and networking scaling policies.
- Engineer, standardize, and maintain our end-to-end CI/CD pipelines, leveraging GitHub Actions in conjunction with ArgoCD and Helm to automate software delivery.
- Establish and enforce Infrastructure as Code paradigms across the engineering portfolio, utilizing Terraform and evaluating the adoption of Crossplane to manage complex cloud resources.
- Partner with 16 or more engineering teams to collaboratively design, implement, and evolve technical standards, tools, and operational processes that respect individual team contexts.
- Lead on-call responsibilities and coordinate incident response efforts, driving thorough root cause analysis for complex, cross-team issues to prevent future recurrence.
- Integrate artificial intelligence tools into daily operational and developmental workflows to boost personal productivity and mentor less experienced colleagues in effective AI-assisted practices.
- Act as the primary SRE representative in strategic planning discussions with engineering leadership, providing data-driven insights into capacity, risk, and infrastructure roadmap decisions.
- Continuously evaluate and optimize system performance, cost, and resilience, ensuring our infrastructure can meet current demands and scale for future growth.
- Foster a culture of reliability by creating clear runbooks, automating alerting pathways, and improving the overall operator experience for development teams.
- Design and implement robust database strategies for PostgreSQL and DocumentDB (MongoDB) deployments on AWS RDS and Aurora, ensuring high availability and data integrity.
- Contribute to the development of internal platforms and self-service tools that empower engineering teams to build and deploy applications with minimal operational friction.
Requirements
- Possess 8 to 10 years of cumulative professional experience as a Senior Site Reliability Engineer, DevOps Engineer, or Infrastructure Engineer, with a proven track record of ownership and impact.
- Demonstrate experience operating and scaling systems within high-growth SaaS environments, including Series A, B, and Series C through D companies, where adaptability and ownership in the face of ambiguity are essential.
- Have extensive, hands-on expertise with Kubernetes orchestration, specifically within EKS environments, including deep knowledge of Helm, ArgoCD for GitOps, and ECS.
- Show advanced proficiency in managing CI/CD pipelines, with a strong background in GitHub Actions and the ability to automate complex deployment workflows.
- Bring practical, production-level experience with relational and document database technologies, specifically Postgres and DocumentDB (MongoDB), deployed on AWS RDS and Aurora.
- Exhibit a strong history of successful cross-functional collaboration, including managing project timelines, setting clear expectations, and effectively partnering with engineers from diverse technical backgrounds.
- Provide evidence of actively using AI tools in your current workflow to enhance coding, debugging, design, and operational tasks, and to mentor others in adopting these practices.
- Hold a Bachelor's degree or equivalent practical experience in a related technical field, such as Computer Science, Engineering, or a similar discipline.
Nice to have
- Demonstrated hands-on experience with Crossplane for infrastructure orchestration.
- Professional experience using Datadog for advanced observability, monitoring, and alerting.
Practical notes
This is a remote position located in the USA. The role offers comprehensive benefits, including health, dental, and vision care, life insurance, mental wellness coverage, fertility and growing family support, flexible time off, paid family and medical leave, bereavement leave, retirement plans, a home office setup allowance, and an annual professional development stipend.