Staff Software Engineer
Job description
About the role
owns the end to end delivery and technical leadership of Tempo, the high performance distributed tracing backend that powers Grafana Cloud Traces. You will translate ambiguous product goals into reliable observability platforms for major customers under strict reliability and latency expectations. This role requires setting technical direction for the hardest problems on the Tempo roadmap while elevating the standard of execution across the distributed team. You will champion architectural decisions for storage and query paths that balance cost, latency, and durability at scale. The position is a remote opportunity physically located in Spain, requiring comfort with asynchronous communication and global collaboration. You will partner closely with product, design, and other engineers to evolve Tempo from a SaaS database into a platform that powers Grafana's next generation of observability products. Your work will directly support the shift from foundational work to product and operational excellence, ensuring the platform can serve agent driven workloads and larger bursty traffic. You will define operational playbooks and review checkpoints that catch performance regressions before they reach production workloads. This role impacts customers by eliminating rough edges, confusing limits, and hidden failure modes so Grafana Cloud Traces "just works" for demanding use cases.
Key facts
What you'll do
- Design ingestion pipelines that normalize high cardinality trace data while meeting strict service level agreements and predictable rollout windows.
- Champion architectural decisions for storage and query paths that balance cost, latency, durability, and long term maintainability.
- Define review checkpoints that catch performance regressions before they reach production workloads and external customer dashboards.
- Coordinate shipping strategies that keep rollout windows predictable, rollback paths clear, and change management transparent to downstream users.
- Forge partner integrations that align Tempo capabilities with external observability ecosystems and third party platforms.
- Champion the evolution of TraceQL metrics so downstream analytics gain consistent mathematical behavior across traces and use cases.
- Architect limit handling and autoscaling patterns that support rapid cell growth without manual intervention or operational bottlenecks.
- Build agent friendly interfaces that simplify assistant driven triage and investigation workflows for developers and site reliability engineers.
- Drive performance targets that sustain high throughput while reducing query latency at scale across globally distributed deployments.
- Define operational playbooks that make complex debugging flows reproducible, auditable, and supportable for internal and external teams.
- Prepare Tempo for agent driven workloads, larger bursty traffic, higher cardinality, and new categories of AI powered workflows such as assistant driven triage and "why is this slow" investigations.
- Contribute to aggressive autoscaling, parameterized rollouts, and aggressive toil reduction as the platform grows from close to 50 cells this year into triple digits.
- Ensure end to end reliability of the trace backend by owning critical design reviews, failure mode analysis, and capacity planning exercises.
- Mentor and elevate the standard of code and design reviews across the team, fostering a culture of transparency, autonomy, and trust.
Requirements
- You must have built large scale distributed systems handling trace workloads in production environments with sustained reliability and performance.
- You need deep experience with Golang and pragmatic judgment about when to simplify versus generalize abstractions for long term maintainability.
- You should understand storage formats, indexing strategies, and query execution for time series and trace data in high cardinality scenarios.
- You must be comfortable designing public APIs that remain stable across multiple major releases and backward compatible integrations.
- You must have experience operating observability platforms serving demanding customers and navigating complex production incidents.
- You must be able to work fully remote from Spain with strong written communication skills for asynchronous collaboration across global teams.
- You should demonstrate ownership of end to end product outcomes, including design, implementation, testing, rollout, and postmortem analysis.
- You must be comfortable making technical decisions that balance tradeoffs between performance, cost, durability, and operational simplicity.
Nice to have
- Exposure to open source project governance and contributor workflows is helpful for engaging with the broader Grafana ecosystem.
Practical notes
This is a full time remote role located in Spain with compensation between $180,000 and $220,000 USD per year. The position expects adherence to the remote work policies and guidelines published Travel requirements are not expected as part of this role. Visa sponsorship information and start date details must be confirmed