Staff Software Engineer
Job description
About the role
Grafana Labs is seeking a Staff Software Engineer
SRE to safeguard the reliability of its highest-value Grafana Cloud customers and their critical database workloads. In this role, you will own the production stability of databases powered by Mimir, Loki, Tempo, and Pyroscope delivered as a SaaS platform across AWS, GCP, and Azure. You will partner directly with embedded product engineering squads to define and evolve per-tenant reliability models and SLOs that reflect real-world customer needs. A core responsibility involves designing and implementing automation that scales reliable operations without increasing manual toil or noise. You will serve as the primary escalation authority and lead incident response for complex, customer-impacting outages, driving post-incident reviews to root cause. The position requires proactive SLO burn analysis to prevent repeat incidents and the continuous refinement of alerting to reduce false escalations. You will also contribute to design documents and code reviews, ensuring that new features are built with production operability and scalability from the outset.
Key facts
What you'll do
Partner closely with product engineering squads using an embedded model to ensure requirements align with reliability constraints.
Own production reliability for high-SLA and complex customer environments running on Mimir, Loki, Tempo, and Pyroscope across major cloud providers.
Design and implement automation to scale our reliability practices and eliminate repetitive manual interventions.
Ensure our customers consistently meet contractual SLO targets through measurement, analysis, and iterative improvement.
Define and evolve per-tenant SLOs and reliability models that accurately represent customer usage and risk profiles.
Proactively reduce SLO burn rates to prevent repeat incidents and improve the stability of long-term service objectives.
Serve as a primary escalation point and maintain on-call responsibilities for relevant production incidents and service disruptions.
Lead customer-impacting incident response, coordinating cross-functional teams and driving post-incident reviews to generate actionable learnings.
Contribute to architectural design docs and participate in rigorous code reviews focused on reliability, scalability, and maintainability.
Influence feature design to ensure production scalability, operability, and a positive customer experience from day one.
Build and maintain automation to eliminate toil, streamline operations, and allow engineers to focus on high-value work.
Improve alert quality and fidelity to reduce noisy escalations and ensure the right signals reach the right engineers at the right time.
Leverage modern AI coding assistants within company-funded budgets and security guidelines to accelerate prototyping, testing, and incident follow-ups.
Access frontier large language models such as GPT-Codex 5/3, Claude Opus 4.6, and Gemini 3 Pro to enhance development and debugging workflows.
Requirements
Candidates must bring 8+ years of total engineering experience, with at least 4 years focused on SRE, CRE, or production engineering roles. A strong preference is given to individuals with formal customer reliability engineering backgrounds. You must possess deep expertise in Kubernetes operations within AWS, GCP, or Azure environments, along with fluency in infrastructure-as-code tools such as Helm, Terraform, or Jsonnet. Demonstrated technical leadership is essential, including the ability to lead complex projects, mentor engineers, and act as a force multiplier across multiple teams and initiatives. Experience designing and operating large-scale distributed systems that serve high-SLA customers is required, along with a track record of making sound decisions in ambiguous and fast-paced situations. You should be comfortable managing concurrent priorities, balancing strategic planning with urgent operational demands, and communicating clearly with both technical and non-technical stakeholders. A pragmatic mindset is critical, as you will be expected to balance rapid iteration with robust testing, thorough code review, and strict quality standards. Since this role involves on-call duties, you must be able to reliably respond to production incidents outside standard working hours and participate in rotating schedules.
Nice to have
Only items explicitly indicated as preferred in the source are listed below.
Practical notes
This is a remote opportunity and we are looking for candidates from the UK, Sweden, Spain or Germany.