Staff Software Engineer
Job description
About the role
You will partner closely with product engineering squads in an embedded model to increase the reliability of Grafana Cloud databases that power our highest value customers. You will own production reliability for complex, high-SLA customer environments across Mimir, Loki, Tempo, and Pyroscope running as SaaS on AWS, GCP, and Azure. You will design and implement automation to scale our reliability practices while defining and evolving per-tenant SLOs and reliability models to meet strict targets. You will proactively reduce SLO burn to prevent repeat incidents and serve as a primary escalation point and on-call for relevant production issues. You will lead customer-impacting incident response and facilitate thorough post-incident reviews to drive lasting improvements. You will contribute to design docs and code reviews to uphold architectural standards and reliability best practices.
Key facts
What you'll do
Partner closely with product engineering squads using an embedded model to align reliability initiatives with product roadmaps and customer expectations.
Own production reliability for high-SLA and complex customer environments, ensuring critical issues are identified and resolved with minimal disruption to customer workloads.
Design and implement automation to scale our reliability practices, reducing manual effort and increasing coverage across multi-tenant scenarios.
Monitor, measure, and report on reliability metrics in near real time to ensure customers meet their defined SLO targets consistently.
Define and evolve per-tenant SLOs and reliability models that reflect diverse customer needs, risk tolerances, and business requirements.
Proactively reduce SLO burn by driving root cause analysis and systemic improvements that prevent repeat incidents from occurring.
Serve as a primary escalation point and on-call for relevant production incidents, providing timely guidance and resolution strategies for critical issues.
Lead customer-impacting incident response and facilitate thorough post-incident reviews to drive actionable follow-ups and track remediation effectiveness.
Contribute to design documents and code reviews to uphold architectural standards and embed reliability best practices into product implementations.
Influence feature design from early stages to ensure production scalability, operability, and resilience are considered throughout the development lifecycle.
Build automation to eliminate toil where needed, enabling engineers to focus on higher-value work and improving overall operational efficiency across teams.
Improve alert quality and reduce noisy escalations by refining thresholds, routing logic, and notification rules to ensure signals are actionable and timely.
Continuously assess tooling and workflows, proposing and implementing enhancements that increase resilience and streamline operational processes for both engineers and customers.
Collaborate with cross-functional teams to balance trade-offs between speed, stability, and innovation while maintaining cloud-native services that meet stringent reliability standards.
Requirements
Bring 8+ years of engineering experience with 4+ years focused on SRE, CRE, or production engineering roles, demonstrating consistent ownership of complex, production-critical systems.
Show strong preference for candidates with formal customer reliability engineering experience and a track record of managing high-SLA services in production environments.
Demonstrate deep experience with Kubernetes operations in AWS, GCP, or Azure, including cluster operations, networking, and storage patterns essential for scalable observability platforms.
Show strong familiarity with infrastructure-as-code tooling such as Helm, Terraform, and Jsonnet to manage repeatable, scalable, and version-controlled deployments.
Have proven ability to work with distributed systems, observability data, and high-throughput logging and metrics pipelines including Mimir, Loki, and Tempo at scale.
Experience designing and implementing automation for reliability, including runbooks, alerting frameworks, and incident response workflows that scale across teams.
Possess a solid understanding of service-level objectives, error budgets, and reliability modeling to guide decision-making, capacity planning, and trade-offs in production.
Demonstrate excellent communication skills to coordinate effectively with product engineers, on-call teams, and customers during incident response, reviews, and reliability initiatives.
Nice to have
Candidates with experience contributing to open-source observability projects and a history of engaging with community-driven development practices.
Background working with AI-assisted development tools and workflows, within security and governance guidelines, to accelerate prototyping, testing, and incident follow-ups.
Experience supporting multi-tenant SaaS platforms at global scale, with a demonstrated ability to manage nuanced customer requirements and regional constraints.
Knowledge of financial services, cloud providers, or highly regulated environments where reliability, auditability, and compliance considerations are critical.
Practical notes
This is a remote opportunity and we are looking for candidates from the UK, Sweden, Spain, Germany, or Ireland.
Hours, travel, visa, or deadlines stated in the source material are not specified beyond the requirement to work remotely from one of the listed countries.