Senior SRE, Managed Gateways
Job description
About the role
You will manage the infrastructure and reliability of the fastest-growing product in our portfolio. This role blends high-level platform engineering with hands-on technical leadership to ensure the stability and scalability of our managed services. You will mentor engineers and set the technical direction for critical gateway infrastructure. The position requires translating complex operational challenges into robust, repeatable platform solutions. You will guide enterprise clients through deployment and configuration of our managed offerings. Your work will directly influence the reliability and performance experienced by customers globally. You will act as a bridge between deep technical implementation and strategic product outcomes. This is an opportunity to shape the foundation of a critical product line within a growing engineering organization.
Key facts
What you'll do
- Lead and mentor a team of engineers focused on maintaining Managed Gateway services, fostering growth and technical excellence.
- Design and manage fault-tolerant systems leveraging Kubernetes, Golang, and multi-cloud environments to ensure high availability.
- Manage the full operational lifecycle of gateway services, including incident response, proactive monitoring, and in-depth post-mortem analysis.
- Build automation and self-service tools that streamline developer workflows and increase overall system efficiency and reliability.
- Establish and monitor SLOs and SLIs, using data to drive decisions that maintain and improve service performance.
- Work directly with enterprise customers to guide them from initial setup through full production implementation and optimization.
- Translate complex customer requirements and diverse deployment patterns into repeatable playbooks and extensible platform features.
- Act as the primary technical escalation point for North American enterprise accounts, providing clear communication and resolution paths.
- Define and enforce platform standards that balance security, compliance, and operational efficiency for managed gateway deployments.
- Analyze system telemetry and logs to identify trends, predict potential failures, and drive architectural improvements proactively.
- Collaborate closely with product and engineering teams to influence the roadmap for gateway features and operational capabilities.
- Develop and maintain detailed runbooks and documentation to ensure operational consistency and ease of knowledge transfer.
- Evaluate emerging technologies and tools to determine their applicability for enhancing our managed gateway infrastructure.
- Partner with security teams to ensure that all gateway services adhere to best practices and regulatory requirements.
Requirements
- Extensive background in Site Reliability Engineering for distributed, high-availability systems with a proven track record of managing complex infrastructures.
- Advanced proficiency with Kubernetes and cloud-native architecture across major platforms such as AWS, GCP, or Azure, including cluster operations and networking.
- Strong coding skills in Golang for building automation, tooling, and custom scripts that enhance operational capabilities.
- Experience managing CI/CD pipelines and Infrastructure as Code using robust tools like Terraform or Ansible to ensure reliable deployments.
- Deep knowledge of observability stacks including Prometheus, Grafana, ELK, or Datadog for monitoring, alerting, and log analysis.
- Ability to manage technical relationships with enterprise clients, communicating effectively with both technical and non-technical stakeholders.
- Demonstrated experience in on-call rotation for critical production systems, with a methodical approach to incident management and resolution.
- A strong understanding of networking fundamentals, including load balancing, TLS termination, and firewall configurations in cloud environments.
Nice to have
- Experience with Service Mesh tools such as Istio or Linkerd for managing microservice communication and resilience.
- Background in database administration for high-throughput systems like PostgreSQL or Cassandra, focusing on performance and scalability.
- History of contributing to open-source SRE projects, demonstrating a commitment to the community and collaborative problem-solving.
- Relevant certifications such as CKA or AWS Certified DevOps Engineer that validate your skills and expertise in cloud operations.
Practical notes
This position is based in Canada and is fully remote. The engagement is full-time, and we encourage you to apply even if you do not meet every listed requirement, as we value candidates who demonstrate strength in specific areas and a willingness to grow.