Staff Backend Engineer - Databases Tempo
Leads the architecture, performance, and operational excellence of Tempo, Grafana’s distributed tracing backend. The role requires substantial distributed-systems experience, strong systems programming ability, production ownership, and technical leadership across multi-quarter initiatives.
About the job
Responsibilities
- Lead multi-quarter technical initiatives from problem framing through rollout, including trace aggregation APIs, Limitless Tempo, autoscaling cells and customer limits, and query engine improvements.
- Own the architecture of core Tempo components: ingestion, storage, query, and metrics generation.
- Drive design reviews, make trade-offs across performance, cost, and complexity, and document architectural decisions.
- Design structured, deterministic, discoverable APIs for humans, agents, downstream products, and external integrators.
- Drive operational excellence against SLOs such as P99 write latency, incident recurrence, and total cost of ownership per ingested GB.
- Automate operations through parameterized rollouts, actionable alerts, and toil reduction.
- Partner with Product and sibling teams to understand how Tempo is consumed and unblock downstream products.
- Mentor engineers through code review, design feedback, pairing, and technical writing.
- Participate in on-call and incident response, including post-incident learning.
- Contribute to the open-source Tempo project and engage with its community.
Requirements
- Track record of leading complex, multi-quarter initiatives spanning design, delivery, and operations.
- Substantial hands-on experience building and operating distributed data systems in production, such as ingestion pipelines, storage engines, or query execution systems.
- Strong software craftsmanship, including clean, robust, maintainable, and performant code.
- Strong Go experience, or substantial experience in Rust, C, or C++ with the ability to learn Go.
- Experience owning production services, participating in on-call, reducing toil, and operating against SLOs.
- Customer-focused, pragmatic approach with iterative delivery and short feedback loops.
- Clear communication through design documents, reviews, and shipped code in a fully remote, asynchronous environment.
Nice-to-haves
- Experience with tracing, OpenTelemetry, or large-scale observability systems.
- Experience designing query languages, SQL- or TraceQL-like engines, or programmatically consumed APIs.
- Experience with columnar storage formats such as Parquet or purpose-built on-disk formats for analytical workloads.
- Experience operating multi-tenant, multi-cell SaaS infrastructure at scale on Kubernetes.
- Experience building for AI/LLM consumers, including structured APIs, metadata and discovery endpoints, deterministic outputs, and evaluation harnesses.
Compensation and benefits
- Company-funded usage budget for modern AI coding assistants.
- Access to frontier models, including GPT-Codex 5/3, Claude Opus 4.6, and Gemini 3 Pro.
Skills
Go, Rust, C, C++, Kubernetes, OpenTelemetry, Traceql, Parquet, Distributed Systems, SaaS, API Design, Query Engines, Autoscaling, SLOs, Open Source
Similar jobs
Backend Engineering jobsBuild and operate Go-based automation and a PostgreSQL-as-a-service platform for high-throughput, always-on production systems. The role requires deep PostgreSQL production experience, backend development expertise, infrastructure automation, and staff-level technical leadership.
Staff Backend Engineer responsible for evolving Grafana into a scalable, multi-tenant observability application platform. The role requires production SaaS experience, distributed-systems expertise, strong backend coding skills, and familiarity with or willingness to learn Golang.
Provides cross-team technical leadership for large backend initiatives, modular architecture, production operations, and AI-assisted engineering adoption. The role requires deep backend architecture experience, monolith modernization expertise, strong written influence, and mentoring ability.
Senior Staff Software Engineer defining the technical vision and architecture for FinHub’s financial ledger and money-movement platform. The role requires 12+ years of backend distributed-systems experience and deep expertise in high-consequence transaction systems.
Build and operate high-performance distributed inference infrastructure serving Claude at global scale, including routing, orchestration, autoscaling, deployment pipelines, and accelerator integration. The role requires significant software engineering experience with distributed systems; experience in large-scale ML infrastructure is preferred.