Senior Engineering Manager, Site Reliability
Job description
Senior Engineering Manager, Site Reliability at Horizon3.ai.
About the role
Horizon3 is actively seeking a strategic leader to establish and scale the Site Reliability Engineering organization from the ground up. This individual will own the design and execution of the reliability strategy that underpins our cybersecurity platform. The successful hire will build and lead a high-caliber SRE team, fostering a culture of operational excellence and continuous improvement. They will be responsible for translating business requirements into robust technical operations standards that ensure system resilience. The role requires balancing hands-on technical oversight with the managerial duties of planning and execution. You will own the end-to-end lifecycle of reliability initiatives, from initial design through implementation and optimization. This position demands a proactive approach to defining processes that prevent issues before they impact customers. Ultimately, you will own the mission of making our platform exceptionally reliable and performant for our users.
Key facts
What you'll do
- Architect and staff a new Site Reliability Engineering department by hiring 4-6 senior engineers within the first twelve months.
- Author and standardize incident management frameworks, establishing clear runbooks for both the SRE group and application engineering units.
- Implement observability best practices by deploying and optimizing APM, logging aggregation, distributed tracing, and metrics collection platforms.
- Cultivate a postmortem culture that emphasizes learning and blameless analysis to drive systemic improvements in platform stability.
- Oversee the reliability roadmap, prioritizing initiatives that reduce toil, improve uptime, and streamline incident response workflows.
- Mentor and coach engineers on-site reliability practices, including on-call rotations, escalation procedures, and diagnostic techniques.
- Partner with product and engineering leadership to define, implement, and monitor service-level objectives and key performance indicators.
- Evaluate and select monitoring and incident response tooling, creating vendor assessment frameworks to guide future technology investments.
- Develop comprehensive process documentation that is clear, accessible, and easily adopted by cross-functional engineering teams.
- Lead training sessions focused on runbook creation, incident simulation exercises, and advanced troubleshooting methodologies.
- Collaborate with infrastructure architects to design scalable and resilient platform solutions using established design patterns.
- Balance the demands of urgent incident response with the strategic execution of long-term reliability and platform enhancement projects.
- Act as the primary liaison between technical operations and executive stakeholders, communicating reliability metrics and program status.
- Champion the adoption of infrastructure as code and architectural decision records to ensure consistency and knowledge retention.
Requirements
- Demonstrated success in leading and expanding SRE or infrastructure teams within a high-growth technology environment.
- Possess prior hands-on experience as a Site Reliability Engineer, with a strong portfolio of implemented solutions.
- Show deep expertise in observability toolchains, including Application Performance Monitoring, centralized logging, distributed tracing, and metric analysis.
- Have direct experience selecting and implementing incident management platforms and constructing vendor evaluation frameworks.
- Bring proven knowledge of at least one major cloud provider, such as Amazon Web Services, Google Cloud Platform, or Microsoft Azure.
- Exhibit strong familiarity with infrastructure or platform engineering design principles, including the creation of architecture decision records.
- Display the capability to author clear, concise, and actionable process documentation that engineering teams will consistently follow.
- Have a history of training engineering staff on runbook maintenance, on-call responsibilities, and structured incident investigation techniques.
- Show aptitude for working collaboratively with product managers and engineering directors to establish measurable SLOs.
- Demonstrate experience in hiring, developing, and retaining engineering talent within remote or hybrid work arrangements.
- Exhibit the skill to manage simultaneous priorities, balancing urgent operational needs against strategic reliability initiatives.
- Communicate effectively with both technical personnel and executive-level stakeholders using appropriate language for each audience.
- Hold the authorization to work in the United States without sponsorship for this position.
- Be willing to commit to occasional travel, up to 10% of time, for business-related engagements.
Nice to have
- Hands-on experience with the NodeZero security validation platform.
Practical notes
- This position is fully remote but requires up to 10% travel for business purposes.
- The compensation package includes competitive salary, equity grants, comprehensive health insurance, vision and dental coverage, flexible vacation policies, and generous parental leave benefits.
- Horizon3 allows candidates to redact age-identifying details, such as graduation dates, from their application materials.
- The work model is full-time and remote, providing flexibility for professionals located within the United States.
- New hires will join a growing engineering organization dedicated to building best-in-class cybersecurity solutions.
- The company is committed to fostering a diverse and inclusive workplace where all employees can thrive.
- Successful candidates must be able to start within a reasonable timeframe as determined by hiring leadership.
- This role reports to the Director of Engineering and works closely with the CTO and cross-functional leadership teams.