Senior Site Reliability Engineer
Job description
About the role
This role bridges software development and operations within a DevOps framework. The position emphasizes production infrastructure reliability and incident management in a global, borderless environment.
Platform and DevOps engineers build the systems that run everything else. They manage infrastructure, CI/CD pipelines, observability, and reliability. The work is about automation, scaling, and removing friction for product teams. Systems thinking is the core skill. Platform teams are measured by developer velocity and system reliability. Most companies run on-call rotations, and understanding incident response is part of the role.
Key facts
What you'll do
Service-level indicators and service-level objectives are defined with teams to establish reliability targets.
Observability systems are built so production behavior remains clear to stakeholders.
Failure scenarios are examined with teams, and possible mitigations are defined to lower risk.
Runbooks are created to address or prevent failure scenarios and guide consistent responses.
Non-value-adding work is reduced to keep systems efficient and focused on outcomes.
Incident management is participated in and facilitated, including on-call duties, to restore service quickly.
Requirements
The posting states a minimum of 2 years of experience.
Five years of experience in software engineering, DevOps engineering, QA engineering and/or cloud engineering are required, including at least two years as a dedicated Site Reliability Engineer.
Assertiveness and strong communication skills are necessary to coach development teams toward reliable choices.
Experience managing incidents in a production environment for a public-facing online service with high business value and high traffic in a 24x7 fashion is required.
Corporate environment experience and navigation of enterprise constraints are expected.
Scripting and programming abilities are used to automate tasks and solve operational problems.
Basic knowledge of serverless services on at least one public cloud provider such as AWS, Azure, or GCP is required.
Familiarity with monitoring systems, including APM tools like Datadog, New Relic, Dynatrace, Prometheus, and Grafana, is required.
Proficiency with pipelining tools such as GitHub, Azure DevOps, GitLab, and Jenkins is required.
Understanding and use of microservices-related technology, including Docker and Kubernetes, is required.
A solid conceptual grasp of software architecture and system thinking is needed to assess design trade-offs.
DevOps context experience and automation for reliability and delivery are required.
Fluency in English at C1 level or above is required for collaboration and incident reporting.
Familiarity with observability platforms such as Datadog is required.
Use of Argo for CI/CD to orchestrate pipelines and deployments is required.
Work with Java and Spring Boot to build and maintain services is required.
Usage of Kafka for event-driven messaging and streaming workflows is required.
Operation of Kubernetes and EKS to manage containerized workloads is required.
Usage of AWS to deploy and run cloud-native services is required.
Experience with publicly accessible, high-availability eCommerce platforms and their reliability demands is required.
Collaboration with on-shore and off-shore teams in international settings is required.
Nice to have
Value stream optimization and reduction of non-value-adding work are considered.
Background in highly regulated or compliance-driven environments is considered.
Practical notes
This position is based in Poland and allows remote work.
The role requires on-call participation and incident management at any hour.
Good to know
Site reliability engineering emphasizes automation, observability, and reliability of distributed systems.
Common tools in this field include monitoring platforms, CI/CD pipelines, container orchestration, and messaging systems.
DevOps practices focus on collaboration between development and operations to enable fast and stable delivery.
Working with public-facing digital services requires attention to availability, resilience, and customer impact.
International teams often operate across time zones and require clear communication and documentation.