Site Reliability Engineering
PatsnapRemote (United Kingdom)Full Time6d ago
RAWSKubernetesAISecuritySaaSLegalGrowthStrategyEngineeringInfrastructurePlatform
Job description
Site Reliability Engineering at Patsnap
About the role
This leadership position involves guiding our Site Reliability Engineering (SRE) team in the UK. You will be instrumental in ensuring our global SaaS platform operates with the highest levels of dependability, security, and efficiency. Your work will directly influence our engineering strategy and the adoption of AI-driven operational practices.
Key facts
What you'll do
- Guide and grow the UK SRE team, setting operational benchmarks, best practices, and reliability targets.
- Maintain the consistent availability, stability, security, and performance of essential platforms and services.
- Shape the operational direction for our worldwide SaaS offering, aiming for superior reliability and performance.
- Manage significant incidents, serving as the primary escalation point during critical system events.
- Define and track key reliability indicators, including service level indicators, objectives, and operational performance metrics.
- Implement automation across infrastructure, deployments, monitoring, and daily operations to boost efficiency and reduce manual tasks.
- Promote the use of AI in operations, employing advanced AI technologies to enhance engineering output and operational quality.
- Collaborate with Engineering, Product, Security, and Infrastructure departments to enhance platform design, scalability, and operational readiness.
- Oversee disaster recovery planning, resilience efforts, and risk management for the entire platform.
- Continuously assess new cloud, AI, and platform technologies to keep Patsnap at the forefront of engineering innovation.
Requirements
- Hold a Bachelor's degree in Computer Science or a comparable discipline, with a minimum of 8 years of experience in DevOps, SRE, or infrastructure management.
- Demonstrate a track record of leading technical teams and overseeing large-scale production environments.
- Possess significant knowledge of cloud environments (AWS preferred), Kubernetes, Docker, CI/CD processes, Infrastructure as Code, and monitoring systems.
- Have a thorough understanding of distributed systems, architectures designed for high availability, and extensive SaaS environments.
- Proven ability to lead initiatives focused on automation and operational excellence.
- Practical experience using AI tools such as ChatGPT, Claude, GitHub Copilot, Codex, or similar to improve engineering efficiency.
- Exhibit strong skills in problem-solving, leadership, communication, and managing relationships with stakeholders.
Nice to have
- Fluency in Mandarin is highly beneficial for effective collaboration with teams in various global locations.
Skills & tools
- AWS
- Kubernetes
- Docker
- CI/CD
- Infrastructure as Code
- Observability Platforms
- AI tools (e.g., ChatGPT, Claude, GitHub Copilot, Codex)
Practical notes
- Patsnap is an equal opportunity employer committed to diversity and inclusion. We encourage applications from all qualified individuals.
- If you do not meet every qualification, we still encourage you to apply and explain your suitability for the role.
- For interview accommodations, please contact recruitment@patsnap.com.