Site Reliability Engineer
Job description
com.
BUILDING THE FUTURE OF OPEN FINANCE
Payward, the parent company behind Kraken, NinjaTrader, Breakout, xStocks, Payward Services, and CF Benchmarks, has invested fifteen years constructing a contemporary financial infrastructure platform designed to promote an open global financial system.
Before you apply, we invite you to review our culture page at https://www.Kraken/culture to understand our driving principles and work methodology.
THE TEAM
Established in 2011, Kraken stands as one of the world's most enduring crypto platforms, serving over 10 million individuals and institutions worldwide. Our offerings include spot trading, margin, futures, staking, and OTC services, all built for individual investors and institutional clients alike.
We are seeking a Site Reliability Engineer to join our Telemetry team. In this role, you will be responsible for helping teams understand, operate, and improve production services. You will work across metrics, logs, traces, alerting, dashboards, and profiling systems to ensure our platform remains observable and scalable.
This is a hands-on position for an engineer who thrives in distributed systems, automation, observability, and production problem-solving. You will collaborate with Software Engineers, Platform teams, Security Engineers, and SREs to enhance operational practices and support mission-critical services. The role includes incident response, on-call responsibilities, platform enhancements, and effective communication of telemetry insights.
The opportunity focuses on owning the reliability and performance of core telemetry systems that power decision-making across the business. You will own the design and evolution of monitoring pipelines that provide clarity into complex, distributed environments. You will own the execution of platform improvements that reduce noise and accelerate signal detection for engineering teams. You will own the automation of deployment and configuration for observability tooling to increase consistency and decrease manual effort. You will own the analysis of incidents to derive durable fixes that prevent recurrence and improve system resilience. You will own the collaboration with cross-functional partners to align telemetry strategy with product and security objectives. You will own the mentorship of best practices in instrumentation and troubleshooting across the organization.
What you'll do
- Operate and enhance the shared platform for metrics, logs, traces, alerting, dashboards, and profiling.
- Manage metrics collection, long-term storage, querying, dashboards, and alerting using Prometheus-compatible systems, VictoriaMetrics, Grafana, and modern alerting tools.
- Oversee log pipelines using Vector, Splunk, and Loki, focusing on reliability, throughput, and troubleshooting.
- Operate distributed tracing and profiling capabilities using Grafana Alloy, Tempo, OpenTelemetry, and Pyroscope.
- Deploy and manage telemetry services using Terraform, Terragrunt, and container orchestration across multiple environments.
- Resolve issues such as missing data, slow queries, broken alerts, pipeline backpressure, and capacity constraints.
- Build reusable configuration and automation to help teams safely manage dashboards, alerts, and telemetry integrations.
- Participate in incident response and on-call, author runbooks, and improve the platform based on real-world learning.
- Partner with software engineers to embed observability into new services from design through production.
- Optimize storage costs and query performance while preserving detailed telemetry for analysis.
- Maintain service level objectives and ensure telemetry accurately reflects system health.
- Conduct postmortem analysis and translate findings into preventative controls and monitoring enhancements.
- Mentor engineers on instrumentation practices and contribute to architectural decisions.
- Track industry trends in observability and propose evaluations for adoption where appropriate.
Requirements
- 3+ years of experience as a Site Reliability Engineer, Platform/Infrastructure Engineer, Observability Engineer, or similar production engineering role.
- Comfortable managing production systems at scale that collect, process, store, and serve telemetry such as metrics, logs, traces, or profiles.
- Experience with Prometheus or a Prometheus-compatible monitoring stack, including metrics collection, querying, and alerting.
- Experience troubleshooting distributed production systems, including availability, latency, data flow, and capacity issues.
- Experience with Infrastructure as Code, particularly Terraform, and CI/CD.
- Experience operating containerised workloads with Nomad, Kubernetes, or similar platforms.
- Solid scripting/programming ability and comfort using AI tools and agents (e.g., Claude) to accelerate delivery.
- Strong incident response, documentation, and collaboration skills.
Nice to have
- Experience with VictoriaMetrics, Grafana, Tempo, Loki, Vector, Splunk, Alertmanager, or OpenTelemetry.
- Experience with PromQL or LogQL, dashboards or alerts as code.
- Experience with maintaining operators and their related CRDs in Kubernetes.
- Experience with Consul, Vault, AWS, and on-premises or datacentre infrastructure.
- Experience operating high-volume logging, streaming, or data pipelines.
- Experience making practical trade-offs between observability data volume, performance, and cost.
- Background in highly regulated or financial services environments where change management and audit trails are critical.
Practical notes
Unless a specific application deadline is stated in the job posting, applications are accepted on an ongoing basis.
Please note, applicants are permitted to redact or remove information on their resume that identifies age, date of birth, or dates of attendance at or graduation from an educational institution.
We consider qualified applicants with criminal histories for employment on our team, assessing candidates in a manner consistent with the requirements of the San Francisco Fair Chance Ordinance.