Head of Site Reliability Engineering
Job description
About the role
Blockchain.com is seeking a Head of Site Reliability Engineering to lead our global infrastructure reliability team based in London. This role oversees the design and implementation of resilient systems that support our cryptocurrency platform and digital asset exchange. You will manage a team of engineers dedicated to ensuring high availability and consistent performance across all production services. The position requires a strategic leader who can balance operational excellence with rapid product development cycles. You will report directly to the executive leadership team to shape the future of our infrastructure roadmap and drive innovation across the organization.
Key facts
What you'll do
1) Lead the site reliability engineering team to maintain platform uptime and system stability across all global regions
2) Architect and implement global infrastructure strategies for the cryptocurrency exchange platform and digital wallet services
3) Drive incident response protocols to minimize downtime during critical system outages and unexpected service disruptions
4) Collaborate closely with product teams to embed reliability practices into the software development lifecycle effectively
5) Manage capacity planning and performance optimization for high-traffic blockchain transaction processing and network validation
6) Establish service level objectives and monitor key metrics across all production and staging environments continuously
7) Oversee the migration of legacy systems to modern cloud-native containerized architectures using Kubernetes orchestration
8) Mentor senior engineers and foster a culture of continuous improvement and technical learning within the group
9) Define comprehensive disaster recovery plans and conduct regular failure injection testing exercises for system resilience
10) Negotiate with cloud providers to secure optimal pricing and resource allocation agreements for infrastructure needs
11) Champion automation initiatives to reduce manual toil and improve deployment frequency and overall system reliability
12) Report directly to the Chief Technology Officer on infrastructure health and reliability trends across the platform
13) Partner with security teams to ensure infrastructure resilience against distributed denial of service attacks
14) Evaluate and adopt new observability and monitoring technologies to enhance system visibility and alerting
Requirements
1) Possess over ten years of experience in site reliability or infrastructure engineering roles within large companies
2) Demonstrated success managing large-scale distributed systems handling millions of daily transactions with high reliability
3) Extensive hands-on experience with Kubernetes orchestration and cloud-native container deployment methodologies in production
4) Strong background in AWS or GCP for building highly available and fault-tolerant cloud architectures
5) Proven track record of leading and growing high-performing engineering teams through challenging technical projects
6) Deep understanding of networking protocols TCP/IP and distributed database systems operating at massive scale
7) Excellent communication skills for translating complex technical concepts to business stakeholders and leadership clearly
8) Ability to thrive in a fast-paced environment while maintaining strict operational discipline and standards
9) Hold a bachelor's degree in computer science or a related technical field from an accredited university
10) Experience with observability tools and establishing comprehensive monitoring dashboards for complex distributed systems
11) Familiarity with agile methodologies and iterative development practices for continuous delivery of reliable software
12) Demonstrated ability to manage multiple concurrent projects while prioritizing critical infrastructure reliability tasks
Nice to have
1) Familiarity with financial regulatory compliance and security standards for digital assets and transactions
2) Prior experience managing on-call rotations and incident management processes at a global scale
3) Contributions to open-source infrastructure projects or community-driven reliability initiatives within the technology industry
4) Knowledge of Terraform or Pulumi for declarative infrastructure provisioning and configuration management workflows
Skills & tools
1) Proficiency in Python or Go for writing automation and operational tooling scripts efficiently
2) Expertise in Prometheus and Grafana for real-time system metrics visualization and alerting
3) Experience with Terraform for managing cloud infrastructure as code configurations and deployments
4) Working knowledge of ELK stack or Splunk for centralized log aggregation and analysis
5) Familiarity with Jenkins or GitLab CI for continuous integration and delivery pipeline setup
6) Understanding of service mesh technologies like Istio for secure microservices communication patterns
Practical notes
1) This role requires a minimum of five years of leadership experience within engineering organizations
2) The position is based in our London office and operates on a full-time schedule
3) Candidates must be authorized to work in the United Kingdom without requiring visa sponsorship
4) The compensation package includes a base salary ranging from one hundred fifty to one hundred eighty thousand pounds annually
5) The application deadline for this position is December thirty-first of the current calendar year