Site Reliability Engineer
Job description
About the role
We are seeking a Site Reliability Engineer to join the Observability group inside our Platform Engineering domain. Platform Engineering's goal is to provide easy to use, self-service platforms to enable other segments to easily build, deploy and monitor their business applications. And Observability's role in that part of the company is to provide our users with end-to-end observability that's easy to use. We are using a modern technology stack to match our principles when it comes to providing the framework for our development team, the company and our customers. In this role, you will build the tools for monitoring and measuring infrastructure, microservices, and sometimes totally unique workloads. You'll put the developer experience at the front of your mind and implementations, and you'll contribute deeply to understanding and preventing incidents through tooling, automation, and people-centered processes. All the while you'll be thinking like an engineer. That means using automation to speed up repetitive tasks and to contribute to the security and compliance of the tech stack that you run and support.
Key facts
What you'll do
Provide end-to-end ownership of observability platforms and pipelines, ensuring they meet evolving business and compliance requirements.
Design, build, and maintain internal self-service tooling that enables developers to instrument, monitor, and troubleshoot their applications independently.
Implement robust data collection by configuring extract mechanisms using Prometheus, StatsD, and OpenTelemetry libraries across diverse environments.
Develop and support resilient data transformation workflows using Vector, Beats, and FluentBit to normalize and enrich telemetry streams.
Construct scalable loading architectures with OpenSearch and Grafana to support high-cardinality metrics, long-term storage, and interactive analysis.
Champion automation practices to eliminate manual effort in monitoring lifecycle tasks, from onboarding new services to alert routing.
Collaborate closely with platform and product teams to define service level objectives, quality gates, and incident response playbooks.
Leverage infrastructure as code with Terraform and AWS CDK to provision and govern cloud resources in a repeatable and auditable manner.
Promote a culture of learning and experimentation, guiding engineers to adopt observability best practices and recover quickly from incidents.
Continuously assess and optimize the cost, reliability, and performance of observability pipelines as data volumes and use cases grow.
Translate complex technical concepts into clear documentation and enablement sessions that help teams adopt tools effectively.
Act as a technical advisor for production incidents, helping to correlate data across metrics, logs, and traces to accelerate root cause analysis.
Maintain and evolve internal dashboards, alerts, and service level indicators to keep them aligned with stakeholder needs and reliability goals.
Explore emerging standards and tools in the observability space, proposing integrations that improve developer experience and operational efficiency.
Requirements
You need to be well-versed in the basics building blocks of observability: metrics, logs, and traces.
You should also be familiar with the tools we need to extract data (think things like Prometheus, StatsD, OpenTelemetry libraries), transform data (tools like Vector, Beats, FluentBit), and load data (OpenSearch, Grafana, Datadog, etc.) so that your fellow engineers can make sense of it.
You must have solid skills with at least one glue language like GoLang or Python to automate workflows and integrate systems.
You need to operate like your brain runs on Linux, with strong comfort and confidence working in command-line environments and shell scripting.
Everything we do is in the cloud, so it is important that you like, understand, and know what it takes to observe cloud-native systems.
You should be comfortable defining and managing your entire environment as code using tools like Terraform and AWS CDK.
You must be able to work autonomously while maintaining a high standard of ownership, communication, and collaboration across distributed teams.
A strong desire to learn and improve, combined with a passion for working with bleeding-edge technologies, is essential for success in this role.
Practical notes
Some roles may require additional in-office presence.
As an N26 employee you will have access to a Premium subscription on your personal N26 bank account. As well as subscriptions for friends and family members.
Additional day of annual leave for each year of service.
A high degree of autonomy and access to cutting edge technologies - all while working with a friendly team of peers of diverse nationalities, life experiences and backgrounds.
Accelerate your career growth by joining one of Europe's most talked about disruptors.
Employee benefits that range from a competitive personal development budget, work from home budget, discounts to fitness & wellness memberships, language apps and public transportation.
Come together with your team in the office for a dedicated day of teamwork each week, plus another day of your choice, and enjoy the flexibility of remote work the rest of the time.