SRE/DevOps
Job description
About the role
The Site Reliability Engineer will own the end to end reliability, performance, and operational excellence of a critical client production environment. You will act as the primary steward of their infrastructure health, designing systems that are observable, automated, and resilient by design. This role requires you to translate abstract reliability goals into concrete, measurable targets and practices that the engineering team can execute against. You will leverage automation and modern AI driven tooling to shift from reactive firefighting to proactive, intelligent operations. The successful candidate will mentor peers, define operational standards, and ensure that reliability is baked into every deployment and architectural decision.
Key facts
What you'll do
Define and enforce measurable reliability standards by establishing SLIs, SLOs, and error budgets that make system behavior transparent and actionable for the entire engineering organization.
Automate all forms of operational toil by converting manual runbooks and procedures into infrastructure as code, self healing mechanisms, and scripted workflows that ensure consistency and speed.
Design, build, and scale the underlying infrastructure required for the client's products, balancing cost efficiency with uncompromising reliability and performance objectives.
Implement and evolve modern CI/CD pipelines to enable frequent, low risk deployments, providing rapid feedback loops and ensuring that releases are safe and predictable.
Partner closely with software engineers and QA teams to embed reliability practices throughout the development lifecycle, from testing strategy to production deployment.
Apply AI and AIOps techniques to automate anomaly detection, suppress alert noise, and enable intelligent remediation that reduces manual intervention.
Lead incident response efforts during outages or degraded states, driving blameless postmortems and converting lessons learned into systemic improvements.
Utilize advanced observability practices with metrics, logs, and distributed tracing to gain deep insight into system behavior and surface latent issues before they impact users.
Mentor engineers across the organization, elevating the overall standard of operational excellence and fostering a culture of shared ownership for production systems.
Continuously evaluate and introduce new tools and techniques, including AI driven platforms, to improve monitoring, reporting, and operational efficiency.
Champion security and compliance considerations within the infrastructure and deployment processes, ensuring that controls align with industry best practices.
Collaborate with cross functional stakeholders to plan capacity, manage changes, and coordinate releases that minimize risk and maximize value delivery.
Maintain detailed documentation of systems, operational procedures, and runbooks to ensure continuity and enable efficient onboarding of new team members.
Act as a technical leader and strategic partner, influencing architectural decisions and long term platform roadmaps to support business objectives.
Requirements
Candidates must possess a minimum of 5 years of hands on experience in Site Reliability Engineering, DevOps, or Platform Engineering roles with a proven ability to own systems from design through operations.
You must demonstrate a deep understanding of reliability engineering fundamentals, including the practical application of SLIs, SLOs, and error budgets to measure and improve system performance.
Hands on experience designing, modernizing, and operating CI/CD pipelines is essential, with fluency in tools that enable safe and efficient software delivery.
Solid expertise in infrastructure as code is required, with proficiency in at least Terraform, Pulumi, or CloudFormation for provisioning and managing cloud resources.
Comprehensive experience with major cloud platforms such as GCP, AWS, or Azure, along with container orchestration platforms like Kubernetes and Docker, is mandatory.
Strong scripting and programming abilities in at least one language such as Python, Go, or Bash are necessary to automate complex operational tasks and integrations.
You must have advanced observability skills, including mastery of metrics, logging, tracing, and alerting using tools such as NewRelic, Prometheus, Grafana, Datadog, or OpenTelemetry.
Applied experience using AI and AIOps tools to automate operations, perform anomaly detection, reduce alert noise, and enable intelligent reporting is a required qualification.
A collaborative mindset is essential, with the ability to work effectively alongside software engineers and QA professionals to achieve shared reliability goals.
Demonstrated technical leadership, including mentoring other engineers, driving cross team initiatives, and influencing engineering practices across the organization, is expected.
Nice to have
Proven experience introducing AIOps tooling into an existing observability stack and integrating it with monitoring platforms.
A background in performance and load testing, including the ability to simulate traffic and identify scalability bottlenecks.
Familiarity with security and compliance frameworks such as SOC 2, ISO 27001, or GDPR, and the ability to implement controls within infrastructure.
Practical notes
Remote Office - Option to work remotely or hybrid.
Parking Space - Free parking available.
Fun Office Space - Game zone and relaxation area.
Health Insurance - Private health insurance, including dental care.
Holidays - 5 extra days after your 1st and 5th year with us.
Personal Development - Company sponsored training and development.
Employee Referral Programme - Competitive bonus for successful referrals.
Social Events - Celebrating success together.
Family Insurance - Add insurance coverage for a family member.
Multisport Card - Fully covered sports pass.