Staff Software Engineer
Job description
About the role
Grafana Labs is seeking a Staff Software Engineer
SRE to partner with high-value Grafana Cloud customers and ensure the reliability of our managed database products built on Mimir, Loki, Tempo, and Pyroscope. In this role, you will own production reliability for complex, large-scale customer environments running on AWS, GCP, and Azure. You will design and implement automation that scales our reliability practices while actively defining and evolving per-tenant SLOs and reliability models. As a primary escalation point and on-call lead, you will lead customer-impacting incident response and drive post-incident reviews to prevent SLO burn and repeat issues. You will contribute to architecture design docs and code reviews while influencing feature design to ensure production scalability and operability across the platform.
Key facts
What you'll do
Partner closely with product engineering squads using an embedded model to align reliability initiatives with product goals.
Own production reliability for high-SLA and complex customer environments across global deployments.
Design and implement automation to scale our reliability practices and reduce manual intervention.
Ensure our customers consistently meet contractual SLO targets through measurement and improvement initiatives.
Define and evolve per-tenant SLOs and reliability models tailored to diverse usage patterns.
Proactively reduce SLO burn to prevent repeat incidents and improve long-term stability.
Serve as a primary escalation point and provide on-call coverage for critical reliability events.
Lead customer-impacting incident response and facilitate thorough post-incident reviews.
Contribute to architecture design docs and participate in rigorous code reviews for reliability-sensitive changes.
Influence feature design to guarantee production scalability, operability, and maintainability.
Build automation to eliminate toil and streamline operational workflows across teams.
Improve alert quality and reduce noisy escalations to focus response on genuine incidents.
Leverage modern AI coding assistants within security guidelines to accelerate prototyping, testing, and incident response.
Access company-funded tools including frontier models such as GPT-Codex 5, Claude Opus 4.6, and Gemini 3 Pro.
Collaborate with cross-functional stakeholders to balance speed, reliability, and long-term technical investments.
Requirements
8+ years of engineering experience with at least 4 years focused on SRE, production engineering, or customer reliability engineering.
Strong Kubernetes expertise across AWS, GCP, or Azure, with deep experience in infrastructure-as-code tools such as Helm, Terraform, and Jsonnet.
Demonstrated history of technical leadership, mentoring engineers, and leading complex projects that scale.
Experience operating distributed systems at high scale with an understanding of performance, resilience, and trade-offs.
Proficiency in scripting and programming to automate operational tasks and build tooling for observability.
Solid understanding of SLOs, error budgets, and reliability modeling for high-availability services.
Proven ability to investigate and resolve complex production incidents in a calm, methodical manner.
Strong written and verbal communication skills for coordinating with product teams and executive stakeholders.
Comfortable working in a fully remote, asynchronous environment with a high degree of ownership and transparency.
Willingness to engage in on-call rotations and respond to customer-impacting issues at any hour when needed.
Nice to have
Formal experience in Customer Reliability Engineering or working with enterprise SRE organizations.
Contributions to open-source observability projects or related tooling in Go, Python, or JavaScript.
Deep knowledge of Grafana, Mimir, Loki, Tempo, or Pyroscope ecosystems.
Experience with cost optimization and performance tuning in multi-cloud environments.
Familiarity with security and compliance best practices for cloud-native platforms.
Practical notes
This role is a remote opportunity available to candidates located in the United Kingdom.
Candidates from Sweden, Spain, or Germany may also be considered based on role requirements.
No compensation details are provided in this source listing.
Work is conducted remotely with an expectation of asynchronous communication and proactive collaboration.
There are no specified travel requirements, visa sponsorships, or deadlines mentioned in this source.