Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 4
Communication @ 6
Grafana @ 4
Kubernetes @ 4
LLM @ 4
Leadership @ 6
Observability @ 4
OpenTelemetry @ 4
Parquet @ 4
Prometheus @ 4
Rust @ 7
SQL @ 4
Security
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Grafana Labs is hiring a Staff Software Engineer to work on Tempo, the open-source distributed tracing backend behind Grafana Cloud Traces and Grafana Enterprise Traces. Tempo enables trace search, span-derived metrics, and connections between tracing data, logs, metrics, and profiles across the Grafana stack.
This role will focus on product and operational excellence, evolving Tempo from a SaaS database into a platform for Grafana observability products, including App Observability, Asserts, Traces Drilldown, and AI-driven assistants.
Responsibilities
- Lead multi-quarter technical initiatives from problem framing through rollout, including trace aggregation APIs, Limitless Tempo, autoscaling cells and customer limits, and query engine improvements.
- Own the architecture of core Tempo components, including ingestion, storage, query, and metrics generation.
- Drive design reviews and make trade-offs involving performance, cost, and complexity.
- Design structured, deterministic, and discoverable APIs for humans, AI assistants, agents, and external integrators.
- Drive operational excellence against SLOs such as P99 write latency, incident recurrence, and total cost of ownership per ingested gigabyte.
- Improve automation, parameterized rollouts, alerting, autoscaling, and toil reduction.
- Partner with Product and sibling teams, including App Observability, Asserts, Drilldown, and Grafana Assistant.
- Mentor engineers through code reviews, design feedback, pairing, and technical writing.
- Participate in on-call for services being built and contribute to incident response and post-incident learning.
- Contribute to the Tempo open-source project, review external contributions, and engage with the community.
- Work on trace aggregation, higher-density APIs, agent-scale ingestion and querying, query performance, multi-cell operations, and customer-facing limits and self-service.
- Use AI coding assistants for prototyping, test generation, refactoring, documentation, and incident follow-ups within security guidelines.
Requirements
- Track record of leading complex, multi-quarter initiatives spanning design, delivery, and operations.
- Substantial hands-on experience building and operating distributed data systems in production, such as ingestion pipelines, storage engines, or query execution systems.
- Strong software craftsmanship, including the ability to write clean, robust, performant, and maintainable software.
- Strong Go experience, or deep experience in another systems language such as Rust, C, or C++ with a path to Go.
- Experience owning production services, participating in on-call, reducing toil, and using SLOs to guide engineering work.
- Customer-focused and pragmatic approach, with the ability to break complex problems into short feedback loops.
- Clear communication and leadership through design documents, reviews, and shipped code in a fully remote, asynchronous environment.
Bonus Qualifications
- Experience with tracing, OpenTelemetry, or large-scale observability systems.
- Experience designing query languages, SQL- or TraceQL-like engines, or APIs consumed programmatically by services or agents.
- Experience with columnar storage formats such as Parquet or purpose-built on-disk formats for analytical workloads.
- Experience operating multi-tenant, multi-cell SaaS infrastructure at scale on Kubernetes.
- Experience building for AI and LLM consumers, including structured APIs, metadata and discovery endpoints, deterministic outputs, and evaluation harnesses.
- Open-source contribution or maintainership experience.
- Experience using Grafana, Prometheus, Loki, or Tempo on call or in a homelab.
- Experience working in a fully remote, globally distributed team.
Working Arrangement
This is a remote opportunity for applicants located in Spain, Sweden, the United Kingdom, Ireland, or Germany. Grafana Labs is a remote-first company that works primarily asynchronously and in writing, with regular video meetings.
Compensation And Benefits
The United Kingdom base compensation range is GBP 103,958–124,750 per year. Actual compensation may vary based on level, experience, and skill set. Benefits include equity, bonus where applicable, and other company benefits. Grafana Labs also provides a remote global culture, career growth pathways, in-person onboarding, and 30 days of annual leave per year, including three Grafana Shutdown Days subject to local legislation.