Staff Software Engineer
Job description
About the role
You will define and drive the technical vision for core Tempo components including ingestion, storage, query, and metrics generation. You will lead multi-quarter initiatives such as trace aggregation APIs, Limitless Tempo, autoscaling cells, and customer limits from problem framing through implementation and rollout. You will set the bar for architecture, performance, and operational excellence across the team by making sharp trade-offs on performance, cost, and complexity. You will design deterministic, structured, and discoverable APIs for both humans and agents to enable Act 3 products and LLM-driven assistants to build reliably on Tempo. You will eliminate rough edges and hidden failure modes so Grafana Cloud Traces "just works" for customers at scale. You will prepare Tempo for an agent-driven world with larger, burstier, higher-cardinality workloads and new AI-powered workflows such as assistant-driven triage. You will own the roadmap from strategy through execution, ensuring that performance targets for query latency and ingestion throughput are met.
Key facts
What you'll do
Define and drive the technical direction for core Tempo components such as ingestion, storage, query, and metrics generation.
Lead multi-quarter technical initiatives from problem framing through rollout, including trace aggregation APIs, Limitless Tempo, autoscaling cells, and customer limits.
Own the architecture of core Tempo components and drive design reviews that make sharp trade-offs on performance, cost, and complexity.
Design APIs for humans and agents, shaping structured, deterministic, and discoverable interfaces for Act 3 products and LLM-driven assistants.
Make Grafana Cloud Traces "just work" for customers by eliminating rough edges, confusing limits, and hidden failure modes.
Achieve operational excellence at scale as the system grows from close to 50 cells today into triple digits this year, focusing on autoscaling, parameterized rollouts, and toil reduction.
Evolve Tempo into a platform enabler with higher-density APIs, trace aggregation, TraceQL metrics math, and machine/LLM-friendly interfaces for downstream products and agents.
Push performance further by delivering faster query latency at hundreds of MB/s ingestion and performant 30-day query ranges to match competitors.
Prepare Tempo for an agent-driven world by supporting larger, burstier, higher-cardinality workloads and new AI-powered workflows such as assistant-driven triage and "why is this slow" investigations.
Collaborate closely with product, operations, and other engineers to ensure that the platform aligns with the broader Grafana Cloud and enterprise observability strategy.
Contribute to open source projects by publishing changes upstream and ensuring that community feedback is incorporated into the roadmap.
Mentor engineers on technical excellence, code quality, and system design to raise the overall bar of the Tempo engineering organization.
Work asynchronously across a fully remote team, communicating clearly through documents, code reviews, and design artifacts.
Continuously assess and mitigate technical debt, scalability risks, and operational hazards to keep the system robust and maintainable.
Requirements
You have a Bachelor's or Master's degree in Computer Science or a related field, or equivalent practical experience.
You have 8+ years of experience in software engineering, including 3+ years in a staff or senior staff role.
You have deep experience designing and operating large-scale distributed systems, including databases, storage engines, and query engines.
You are proficient in systems programming in languages such as Go or Rust, with a strong grasp of performance, concurrency, and networking.
You have a track record of building and operating production services that handle high throughput and low latency at scale.
You understand distributed systems concepts such as consensus, replication, partitioning, and fault tolerance, and can apply them to practical trade-offs.
You have experience with observability concepts such as traces, metrics, and logs, and have worked with telemetry data at scale.
You are comfortable making technical decisions with incomplete information and communicating rationale clearly to both technical and non-technical stakeholders.
You can break down complex problems into manageable milestones and deliver incremental value while maintaining a long-term vision.
You are comfortable reading and contributing to open source projects and engaging with community discussions.
You are able to work asynchronously across time zones and collaborate effectively with a fully remote team.
You care deeply about reliability, performance, and operational simplicity, and you bring a bias for action.
You are passionate about building platforms that enable other teams and products to succeed.
You are based in, or eligible to work in, one of the following countries: Spain, Sweden, UK, Ireland, or Germany.
Nice to have
Experience contributing to open source observability projects such as Tempo, Prometheus, Loki, or similar CNCF projects.
Experience building or operating SaaS platforms at scale with multi-tenancy, billing, and operational tooling.
Familiarity with trace analytics, metrics extraction from traces, and query languages specialized for telemetry.
Knowledge of agent-driven architectures and patterns for AI-assisted observability.
Experience with infrastructure cost optimization and efficiency at large scale.
Practical notes
This role is based in the United Kingdom and is open to applicants located in Spain, Sweden, UK, Ireland, or Germany.
No application page or deadline is provided in the source.