Senior Site Reliability Engineer
Job description
Senior Site Reliability Engineer at Chainalysis Careers.
About the role
The role centers on owning reliability as a core product capability for the Hexagate team at Chainalysis. You will define and evolve service level objectives, alerting standards, and incident response practices to ensure resilient real-time on-chain detection and response. The position requires deep collaboration with backend and platform engineers to reduce toil and prevent recurring failure classes across cybersecurity workloads. You will design and operate observable systems that support high-velocity ingestion, detection, and response pipelines. Driving modernization of infrastructure, including Kubernetes and cloud-native patterns, will be a primary responsibility. Success means enabling other engineers to own services from development through production with durable operational standards. You will take part in root cause analysis and systemic fixes that improve the platform long term. The role is embedded in a fast-moving environment where adversarial threats demand constant vigilance and rapid iteration.
Key facts
What you'll do
- Own reliability as a product capability by defining and evolving SLOs, alerting standards, incident response practices, and production readiness expectations across the platform.
- Build and improve platform foundations that help engineering teams ship safely and quickly, including CI/CD, deployment workflows, and infrastructure automation.
- Design and operate resilient, observable systems that support real-time ingestion, detection, and response workloads across distributed environments.
- Lead improvements in scalability, performance, and operational maturity across services, environments, and deployment pipelines.
- Drive modernization of the infrastructure stack, including Kubernetes-based workloads, infrastructure as code, and standardized operational patterns.
- Improve developer experience by creating internal tooling, paved roads, and clear operational standards so engineers can own services from development through production.
- Partner closely with backend and platform engineers to reduce toil, prevent recurring classes of failure, and raise the reliability bar across the organization.
- Take part in incident management, root cause analysis, and follow-through on systemic fixes - not just short-term mitigation.
- Implement observability strategies that provide deep insight into system behavior, latency, and error modes for complex on-chain data flows.
- Champion operational best practices that align with security and compliance requirements in a cybersecurity-focused product environment.
- Collaborate on capacity planning and infrastructure scaling to meet demands of high-throughput blockchain data processing.
- Refine deployment and rollback strategies to ensure safe, rapid, and predictable releases to production.
- Contribute to the evolution of internal platforms that abstract complexity and enable faster feature development.
- Measure and report on reliability metrics to guide investment and prioritize improvements across the engineering organization.
- Enable cross-team collaboration to standardize patterns that improve stability and reduce operational risk.
Requirements
- 5+ years of experience in SRE, infrastructure engineering, platform engineering, or a closely related role with strong hands-on production ownership.
- Strong experience building and operating cloud-native systems in production, ideally in high-growth SaaS, fintech, cybersecurity, or similarly demanding environments.
- Deep practical experience with Kubernetes, AWS, and infrastructure as code tools such as Terraform or Pulumi.
- Strong understanding of observability, monitoring, debugging, and performance tuning for distributed systems.
- Experience building or maintaining CI/CD systems, deployment tooling, and operational automation.
- Solid coding ability in Python, Go, or Rust, with the ability to build internal tools and automate operational workflows.
- Sound judgment around reliability trade-offs, failure modes, and safe delivery practices.
- A collaborative mindset: you enjoy enabling product teams, mentoring engineers on operational best practices, and building a culture of shared ownership rather than acting as a ticket-based support function.
- High ownership and a bias toward durable solutions that make systems more stable, secure, and scalable over time.
- Comfort working in an environment where threats are constant and response time is critical.
- Willingness to engage directly with production issues and assume responsibility for platform health at all hours when needed.
- Commitment to following and improving operational runbooks, incident processes, and change management practices.
- Understanding of the unique challenges of blockchain and on-chain data environments, including adversarial actors and evolving attack vectors.
Nice to have
- Web3 or blockchain experience, or strong interest in the space.
- Experience with real-time or streaming systems.
- Background in security, fraud, risk, or other adversarial domains.
- Experience building internal developer platforms or defining reliability standards used across multiple teams.
Practical notes
- Full-time position based in the Tel Aviv Office.
- Travel requirements, visa, or application deadlines are not specified in this posting.
Why Chainalysis
Chainalysis provides the tools and data necessary to investigate and understand blockchain activity for the world's most trusted blockchain analytics platform. The environment demands real problem solving with real attackers, creating impact that is immediate and measurable. The company operates as a small team with high ownership, backed by the resources and data advantage of a leading blockchain intelligence firm. AI is integrated into how work gets done, enabling faster responses to threats and new ways of turning instructions into action. The role is suited for individuals who want to shape the future of blockchain intelligence while working at the pace of a startup within a well-established organization.