Staff Software Engineer
Job description
About the role
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand, and act on all their disparate data to move at the speed of their ambitions. Today, more than 35 million users and 7,000+ customers - including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce - trust Grafana Labs to ensure reliability of their applications and systems, resolve incidents quickly, and optimize their telemetry to reduce noise and cost. We are a 100% remote company with 1,600+ team members across 40+ countries, and we're backed by leading investors including Lightspeed Venture Partners, Sequoia Capital, GIC, Coatue, J.P. Morgan, CapitalG, and Lead Edge Capital. Learn more at grafana.com and follow us on LinkedIn and X.
We're scaling fast and staying true to what makes us different: an open-source legacy, a global collaborative culture, and a passion for meaningful work. Our team thrives in an innovation-driven environment where transparency, autonomy, and trust fuel everything we do.
You may not meet every requirement, and that's okay. If this role excites you, we'd love you to raise your hand for what could be a truly career-defining opportunity.
This is a remote opportunity and we are looking for candidates from the UK, Sweden, Spain or Germany.
Key facts
What you'll do
Partner closely with product engineering squads using an embedded model to support our highest value Grafana Cloud customers.
Own production reliability for high-SLA and complex customer environments running databases such as Mimir, Loki, Tempo, and Pyroscope delivered as a SaaS product across AWS, GCP, and Azure.
Design and implement automation to scale our reliability practices and ensure consistent execution at enterprise scale.
Ensure our customers meet their service level objectives by actively managing and driving SLO attainment.
Define and evolve per-tenant SLOs and reliability models that reflect real-world usage and risk.
Proactively reduce SLO burn to prevent repeat incidents and improve long-term system stability.
Serve as a primary escalation point and maintain on-call responsibilities for relevant production incidents requiring immediate attention.
Lead customer-impacting incident response and facilitate thorough post-incident reviews to drive continuous improvement.
Contribute to design docs and participate in code reviews to uphold engineering standards and reliability best practices.
Influence feature design to ensure production scalability and operability are considered from the earliest stages of development.
Build and maintain automation to eliminate toil and improve the efficiency of routine operational tasks.
Improve alert quality and reduce noisy escalations to ensure signals remain actionable and trustworthy.
Requirements
8+ years of engineering experience with 4+ years focused on SRE, CRE, or production engineering in customer-facing reliability roles.
Strong Kubernetes experience across AWS, GCP, or Azure, along with proficiency in infrastructure-as-code tooling such as Helm, Terraform, and Jsonnet.
Demonstrated technical leadership through leading projects, mentoring engineers, and acting as a force multiplier across multiple teams.
Experience managing high-SLA services in distributed, multi-cloud environments that require rigorous reliability practices.
Solid understanding of observability concepts, metrics, logs, and traces as they apply to large-scale SaaS platforms.
Comfort with scripting and automation to operate complex systems at scale and reduce manual intervention.
Excellent written and verbal communication skills for coordinating with product teams and responding to incident communications.
Ability to work asynchronously within a fully remote, globally distributed team spread across multiple time zones.
Nice to have
Formal experience in customer reliability engineering or working within a high-SLA SaaS environment.
Contributions to open source observability projects including Mimir, Loki, Tempo, or Pyroscope.
Experience with AI coding assistants and familiarity with modern developer tooling that supports rapid iteration.
Practical notes
This role is a full-time position (40 hours per week).
We are looking for candidates located in the UK, Sweden, Spain, or Germany.
No travel is required for this role.
No visa sponsorship is provided for this position.
Applications will be reviewed on a rolling basis until the role is filled.