Senior Software Engineering Manager, Managed Gateways SREs
Job description
About the role
This position involves building and leading a new Site Reliability Engineering team based in Toronto to support Kong's managed services. You will balance management responsibilities with direct technical contributions to ensure high availability and performance for global enterprise customers. The role requires you to act as both a hands-on engineer and a people leader who removes blockers and drives operational excellence. You will translate business requirements into robust technical solutions that keep critical services online and performant. Collaboration with product, security, and infrastructure teams will be essential to delivering reliable experiences. Your work will directly influence the reliability posture of Kong's managed offerings across the Americas and Europe.
Key facts
What you'll do
- Establish the Toronto SRE team by recruiting talent and setting operational standards that define how reliability is executed.
- Participate directly in technical implementations and reliability projects during the team formation phase to ensure delivery against demanding timelines.
- Mentor engineers to maintain high performance and service stability for clients in the Americas and Europe through coaching and technical guidance.
- Design and manage scalable, fault-tolerant systems using Kubernetes and cloud infrastructure to meet enterprise availability targets.
- Manage the full operational lifecycle, including incident response, monitoring, and post-mortem analysis to extract learning and drive improvements.
- Develop automation and self-service tools to improve deployment workflows and reduce manual labor, enabling faster and safer changes.
- Establish and monitor Service Level Objectives and Indicators to guide architectural improvements and capacity planning decisions.
- Work with Product and engineering departments to align operational readiness with feature development, ensuring services are reliable by design.
- Define and evolve the technical roadmap for the Managed Gateways SRE team in Toronto, balancing quick wins with strategic initiatives.
- Act as the operational voice of the customer within engineering discussions, ensuring reliability constraints are considered early and often.
- Build and maintain runbooks and documentation that enable consistent execution and smooth onboarding of new team members.
- Partner with security and compliance stakeholders to ensure that reliability practices meet regulatory and organizational standards.
- Leverage observability pipelines to detect issues before they impact customers and to provide deep insights during investigations.
- Foster a culture of blameless postmortems and continuous improvement where data drives decisions and process changes.
Requirements
- Demonstrated history of managing SRE or DevOps teams within high-growth organizations where reliability practices matured quickly.
- Technical expertise in operating distributed systems on AWS, Azure, or GCP, with evidence of production workloads at scale.
- Production-level experience with Kubernetes and container orchestration, including cluster lifecycle management and networking.
- Programming proficiency in Golang for automation and service development, with a portfolio of internal tools or libraries.
- Knowledge of observability practices using tools like Prometheus, Grafana, or OpenTelemetry to measure and analyze system behavior.
- Experience managing critical system incidents and performing root cause analysis to prevent recurrence and improve resilience.
- Strong understanding of networking, load balancing, and security fundamentals as they apply to distributed services and gateways.
- Ability to balance strategic planning with hands-on execution, contributing code while guiding technical direction.
Nice to have
- Practical knowledge of API gateways, service mesh, or network proxies that aligns with Kong's product domain and technology stack.
- Active involvement in open-source projects or cloud-native communities that demonstrate engagement with the broader ecosystem.
- Professional certifications such as CKA, CKAD, or AWS Certified DevOps Engineer that validate skills in orchestration and cloud operations.
- Previous work experience in infrastructure software, developer tools, or API management sectors that provide context for platform product decisions.
- Experience mentoring engineers in fast-paced environments where growth and learning are constant priorities.
- Familiarity with cost optimization strategies for cloud infrastructure that align reliability goals with business outcomes.
Practical notes
Candidates who do not meet every listed requirement are encouraged to apply, as the hiring team values a mix of strengths across different areas. The compensation range reflects the market rate for this role in Toronto and includes a comprehensive benefits package. This position is full-time and based in Toronto, with potential for occasional travel within the Americas and Europe as needed for collaboration or customer engagements. There are no specific visa sponsorship details listed at this time, and interested candidates should highlight how their experience aligns with the outlined responsibilities during the application process.