Staff Software Engineer
Job description
About the role
Grafana Labs is seeking a Staff Software Engineer
SRE to own the production reliability of our highest value Grafana Cloud customers running Mimir, Loki, Tempo, and Pyroscope. The role centers on ensuring that cloud databases delivered as a SaaS product across AWS, GCP, and Azure meet strict service-level agreements in complex, real-world environments. You will partner closely with embedded product engineering squads to drive reliability initiatives, design and implement automation that scales reliability practices, and proactively reduce SLO burn to prevent repeat incidents. This position requires serving as a primary escalation point, leading customer-impacting incident response, and owning post-incident reviews. You will define and evolve per-tenant SLOs and reliability models that align with actual usage patterns and risk profiles. The role demands strong ownership of production systems and influence over feature design to ensure operability at scale, while contributing directly to architecture design documents and code reviews to uphold long-term system integrity. This is an opportunity to shape reliability practices for a global observability platform used by major enterprises.
What you'll do
Partner with product engineering squads using an embedded model to drive reliability initiatives and align technical strategies.
Own production reliability for high-SLA and complex customer environments across AWS, GCP, and Azure, ensuring databases remain performant and available.
Design and implement automation that scales reliability practices and reduces manual intervention, enabling teams to operate at speed.
Ensure Grafana Cloud databases meet defined SLO targets under demanding conditions and evolving customer demands.
Define and evolve per-tenant SLOs and reliability models to reflect actual usage, risk tolerance, and operational constraints.
Proactively reduce SLO burn to prevent repeat incidents and improve long-term stability through predictive and reactive measures.
Serve as a primary escalation point and provide on-call support for critical database incidents that impact customer experience.
Lead customer-impacting incident response and facilitate thorough post-incident reviews to drive continuous improvement.
Contribute to architecture design documents and participate in rigorous code reviews to uphold reliability and maintainability standards.
Influence feature design to guarantee production scalability and operability from the start, balancing innovation with stability.
Build and maintain automation to eliminate repetitive toil and improve engineering efficiency across reliability workflows.
Improve alert quality and reduce noisy escalations through better detection, routing, and prioritization logic.
Leverage modern AI coding assistants within security guidelines to accelerate prototyping, test generation, incident response, and documentation.
Collaborate cross-functionally to align reliability goals with product roadmaps, customer expectations, and business outcomes.
Requirements
Candidates must bring 8+ years of broad engineering experience, with at least 4 years focused on SRE, CRE, or production engineering in customer-facing, high-availability systems.
Strong preference is given to individuals with formal customer reliability engineering experience and a proven track record in operating critical services.
Demonstrate deep expertise with Kubernetes across major cloud providers, including AWS, GCP, and Azure, and hands-on experience managing distributed databases at scale.
Proficiency with infrastructure-as-code tooling such as Helm, Terraform, and Jsonnet is essential for defining, deploying, and managing reliable environments.
Experience leading technical initiatives and mentoring engineers to multiply team effectiveness, including guiding less experienced staff through complex reliability challenges.
Ability to translate complex customer requirements and production constraints into robust reliability models, automation strategies, and actionable runbooks.
Strong written and verbal communication skills for coordinating with product teams, incident responders, and executive stakeholders to ensure alignment and transparency.
Comfortable operating in a fast-paced, remote-first environment with high ownership, minimal supervision, and a bias for action in the face of ambiguity.
Willingness to rotate on-call duties and respond promptly to critical production issues at any time, including diagnosing root causes and coordinating mitigations.
Capability to work independently and drive projects forward without dependency bottlenecks, taking full ownership of reliability outcomes and follow-through.
Commitment to maintaining detailed design documentation, runbooks, and operational playbooks that evolve with system complexity and incident learnings.
Openness to using AI-assisted development tools while adhering to strict security, compliance, and quality standards, ensuring generated artifacts are reviewed and integrated responsibly.
Passion for observability platforms and improving the reliability of large-scale SaaS services, with a focus on customer trust and operational excellence.
Practical notes
This role is a remote opportunity open to candidates from the UK, Sweden, Spain, or Germany, reflecting regional employment and compliance considerations.
No compensation details are provided in the source text, and candidates are encouraged to discuss expectations during the hiring process if required.
No travel, visa, or application deadline information is specified in the source materials; interested candidates are advised to check official channels for current opportunities.