Staff Backend Engineer - Databases Tempo | Canada | Remote

📍 Canada
CAD 186,400-223,600 per year
SENIOR
✅ Remote

Tech Stack

AI @ 4 API @ 4 Communication @ 7 Grafana @ 4 Kubernetes @ 4 LLM @ 4 Observability @ 4 OpenTelemetry @ 4 Parquet @ 4 Prometheus @ 4 Rust @ 7 SQL @ 4 Security

Details

Grafana Labs is hiring a Staff Backend Engineer to work on Tempo, its open-source distributed tracing backend behind Grafana Cloud Traces and Grafana Enterprise Traces. This remote role is open to candidates in Canada and the United States.

Tempo is evolving from a SaaS database into a platform supporting Grafana observability products, including App Observability, Asserts, Traces Drilldown, and AI-driven assistants. The role focuses on product and operational excellence, scalability, performance, agent-friendly APIs, and large-scale distributed data systems.

Responsibilities

  • Lead multi-quarter technical initiatives from problem framing through rollout, including trace aggregation APIs, Limitless Tempo, autoscaling cells, customer limits, and query engine improvements.
  • Own the architecture of core Tempo components, including ingestion, storage, query, and metrics generation.
  • Drive design reviews and make trade-offs involving performance, cost, and complexity.
  • Design structured, deterministic, and discoverable APIs for Grafana products, LLM-driven assistants, and external integrators.
  • Drive operational excellence against SLOs such as P99 write latency, incident recurrence, and total cost of ownership per ingested gigabyte.
  • Improve automation, parameterized rollouts, actionable alerts, and Zero Ops practices.
  • Partner with Product and sibling teams, including App Observability, Asserts, Drilldown, and Grafana Assistant.
  • Mentor engineers through code reviews, design feedback, pairing, and technical writing.
  • Participate in on-call for services developed by the team and contribute to incident response and post-incident learning.
  • Contribute to the Tempo open-source project, review external contributions, and engage with the community.
  • Work with modern AI coding assistants for prototyping, test generation, refactoring, documentation, and incident follow-ups, within security guidelines.

Example Projects

  • Extending TraceQL metrics and designing LLM-friendly response types.
  • Building end-to-end autoscaling with customer limits, Tempo cells, hysteresis, predictive scaling, and safe scale-down.
  • Adding guardrails for bursty, high-cardinality, agent-generated workloads.
  • Improving query performance through new data formats, smarter query pipelines, and targeted optimizations for Drilldown and Traces workflows.
  • Supporting 30-day query ranges.
  • Building parameterized rollouts and multi-cell operations tooling.
  • Improving customer-facing configuration and observability for limits and self-service.

Requirements

  • A track record of leading complex, multi-quarter initiatives spanning design, delivery, and operations.
  • Deep hands-on experience building and operating distributed data systems in production, such as ingestion pipelines, storage engines, or query execution systems.
  • Strong software craftsmanship and experience writing clean, robust, performant, maintainable software.
  • Strong Go experience, or substantial experience with systems languages such as Rust, C, or C++ and a path to learning Go.
  • Experience owning production services, participating in on-call, reducing operational toil, and treating SLOs as a product feature.
  • A customer-focused, pragmatic approach based on short feedback loops, MVP delivery, learning, and iteration.
  • Strong written and verbal communication skills, with experience leading through design documents, reviews, and shipped code in a fully remote, asynchronous environment.

Bonus Qualifications

  • Experience with tracing, OpenTelemetry, or large-scale observability systems.
  • Experience designing query languages, SQL- or TraceQL-like engines, or programmatically consumed APIs.
  • Experience with columnar storage formats such as Parquet or purpose-built on-disk formats for analytical workloads.
  • Experience operating multi-tenant, multi-cell SaaS infrastructure at scale on Kubernetes.
  • Experience building structured APIs, metadata or discovery endpoints, deterministic outputs, or evaluation harnesses for AI and LLM consumers.
  • Open-source contribution or maintainership experience.
  • Experience using Grafana, Prometheus, Loki, or Tempo on call or in a homelab.
  • Experience working in a fully remote, globally distributed team.

Work Environment

Grafana Labs is a remote-first, remote-only company. The team works asynchronously and communicates primarily in writing, with regular video meetings. The company provides access to AI coding assistants and a company-funded usage budget. The role includes on-call participation and an in-person onboarding program.

Benefits

  • Compensation in Canada: $186,368–$223,642 CAD per year.
  • Restricted Stock Units (RSUs) are included with all roles.
  • Global annual leave policy of 30 days per year, including three Grafana Shutdown Days, subject to local legislation.
  • Remote global culture, career growth pathways, transparent communication, and empowered teams.
  • Equal-opportunity employment practices.

More jobs at Grafana Labs

Similar jobs