Site Reliability Engineer
Job description
About the role
The role shapes infrastructure evolution for a global financial platform serving millions. The SRE collaborates with product and engineering teams to scale distributed systems securely and reliably. Core responsibilities focus on reliability, observability, and efficient delivery.
Platform and DevOps engineers build the systems that run everything else. They manage infrastructure, CI/CD pipelines, observability, and reliability. The work is about automation, scaling, and removing friction for product teams. Systems thinking is the core skill. Platform teams are measured by developer velocity and system reliability. Most companies run on-call rotations, and understanding incident response is part of the role.
Key facts
What you'll do
Infrastructure evolves to address reliability, latency, bandwidth, and security with measurable outcomes for millions of customers.
Observability, monitoring, and alerting deliver clearer system visibility and faster response across the platform.
Engineering teams coordinate work to define the most efficient execution path and reduce bottlenecks.
Common workflows are centralized where possible, replacing duplicated efforts with shared services and standard patterns.
Tooling replaces manual, repetitive tasks, enabling scalable operations and freeing teams for higher-value work.
Engineers complement a high-caliber team in a fast-paced, dynamic environment, maintaining high standards.
Requirements
Containerization and service orchestration follow best practices, with security integrated throughout, including experience using Hashicorp Nomad, Consul, and Vault.
Proficiency in at least one programming language is required, with experience in Golang, Python, and Bash as a plus.
Linux knowledge includes resource allocation, networking, and internals for effective system management.
Experience with cloud solutions on GCP or AWS supports platform scalability and resilience.
Modern monitoring tools such as Prometheus, Datadog, Grafana, and Telegraf are used to maintain visibility and performance.
Infrastructure as code practices are applied, with complex Terraform deployments preferred for consistency and automation.
Configuration management relies on tools such as Saltstack to maintain stable, repeatable environments.
GitOps and CI pipelines, especially using GitHub Actions, enable safe and efficient changes.
Messaging systems like Kafka move data reliably across distributed services.
Database management ensures data integrity, scalability, and performance across storage technologies.
Experience in data centers informs capacity planning and operational decisions where relevant.
Routing and switching protocols knowledge supports network design and troubleshooting in complex environments.
Practical notes
Blockchain is an equal opportunity employer; decisions depend solely on qualifications, merit, and business needs.
Personal data you provide is processed to manage recruitment, conduct interviews and tests, and evaluate applications in line with applicable regulations. Typical interview steps
Platform interviews usually include an infrastructure scenario, a scripting or coding exercise, and operational questions. Candidates may be asked to design a deployment pipeline or debug an outage. Incident experience and an automation mindset are tested. Interviewers often ask about a past outage and how you handled it. Structured post-incident thinking, not heroics, is what they look for.
Nice to have
Experience working in data centers is a plus.
Skills & tools
Proficiency with container orchestration, monitoring platforms, configuration management, and CI/CD pipelines.
Good to know
The role focuses on platform reliability, observability, and automation at scale.
Common tools include container orchestration systems, monitoring platforms, and infrastructure as code solutions.
The team values abstract thinking, rapid delivery, and minimizing low-effort operational work.
Engineers are expected to propose, discuss, and implement infrastructure changes autonomously.
This position operates in a regulated financial technology environment with strict security expectations.
Career growth
Platform careers grow from engineer to senior, staff, and platform lead roles. Some people move into SRE leadership or cloud architecture. Breadth across networking, storage, and reliability becomes more important at senior levels. Platform careers reward breadth and calm under pressure. Experience automating your own work is the strongest signal for senior roles.