Senior Manager, Site Reliability Engineering
Job description
About the role
We are seeking a technical leader to oversee our Site Reliability Engineering initiatives. You will guide a team focused on maintaining high availability and performance for our identity infrastructure as we expand our support for AI-driven security. In this capacity, you will be responsible for ensuring that the foundational platforms supporting our customers remain resilient and performant under demanding conditions. The role requires a strategic mindset that can align technical execution with broader business objectives in a rapidly evolving security landscape. You will act as a bridge between operational realities and architectural vision, translating reliability requirements into actionable engineering initiatives. This position is critical for maintaining the trust of our customers who depend on our systems for their most sensitive identity workflows. Your leadership will directly influence the stability and scalability of the services that power modern digital interactions.
Key facts
What you'll do
- Direct the day-to-day management and professional growth of a team of engineers dedicated to system reliability and uptime, fostering an environment of continuous learning and improvement.
- Architect and oversee the long-term strategy for our large-scale identity platforms, ensuring they meet current demands and are adaptable to future technological shifts.
- Partner closely with cross-functional engineering groups to refine deployment pipelines, enhance system observability, and integrate best practices across the organization.
- Lead the development and execution of incident response protocols, driving thorough post-mortem analysis to identify root causes and implement preventative measures.
- Evaluate and balance immediate operational demands with the necessity of building scalable, robust infrastructure that supports sustainable growth.
- Champion the adoption of Site Reliability Engineering principles to automate operations and improve the efficiency of our reliability workflows.
- Provide technical guidance on distributed systems architecture, ensuring that our identity infrastructure is designed for resilience, security, and optimal performance.
- Utilize advanced cloud infrastructure management techniques to optimize resource utilization, control costs, and maintain high standards of service delivery.
- Leverage sophisticated incident management and system observability tools to monitor platform health, detect anomalies, and respond to issues proactively.
- Serve as a subject matter expert and thought leader within the organization, contributing to the development of standards and methodologies for reliability engineering.
- Facilitate clear communication between technical and non-technical stakeholders to ensure alignment on reliability goals and the status of critical initiatives.
- Drive the implementation of monitoring and alerting frameworks that provide deep insights into system behavior and potential vulnerabilities.
- Guide the team in conducting blameless post-mortems that generate actionable insights and drive meaningful changes to prevent future disruptions.
- Evaluate emerging technologies and industry trends to determine their applicability in enhancing the reliability and security of our identity platforms.
Requirements
- Bring proven experience in a management or leadership role within an SRE or DevOps environment, demonstrating the ability to lead through influence and technical acumen.
- Possess a strong background in distributed systems and cloud infrastructure, with a clear understanding of how these technologies operate at scale.
- Show the capability to operate with urgency and composure in high-stakes production environments where quick decision-making is essential.
- Have a demonstrated history of successfully building, mentoring, and scaling engineering teams to achieve ambitious technical and business goals.
- Exhibit a deep commitment to the Site Reliability Engineering discipline, balancing automation, performance, and reliability to support business outcomes.
- Hold a solid understanding of identity and access management concepts to effectively support the specific needs of our customer base.
- Thrive in a collaborative setting, working effectively with diverse teams to solve complex problems and deliver high-quality solutions.
- Maintain strong written and verbal communication skills to articulate technical concepts to both technical and non-technical audiences.
Nice to have
The source listing does not specify any preferred skills, experiences, or qualifications beyond the core requirements.
Practical notes
The engagement is Full-time, indicating a standard 40-hour work schedule. The position is based in Washington, DC, which may require adherence to specific local working hours. The role does not mention international travel, visa sponsorship, or specific deadlines for application submission in the provided source material.