Staff Software Engineer - Databases SRE

📍 Germany
📍 Spain
📍 United Kingdom
📍 Sweden
EUR 109,700-131,700 per year
SENIOR
✅ Remote

Tech Stack

AI @ 7 AWS @ 7 Azure @ 7 Design Patterns GCP @ 7 Go @ 4 Grafana @ 9 Helm @ 3 Java @ 4 Kubernetes @ 7 Leadership @ 7 Linux @ 4 Mentoring @ 7 Networking @ 4 Observability Python @ 4 SRE @ 9 Security @ 7 Technical Leadership @ 7 Terraform @ 3

Details

Grafana Labs is seeking a Staff Software Engineer - SRE to support its highest-value Grafana Cloud customers by improving the reliability of cloud databases based on Mimir, Loki, Tempo, and Pyroscope. These databases are provided as a SaaS product across AWS, GCP, and Azure regions.

The SRE team is embedded within the Mimir, Loki, and Tempo squads and is responsible for ensuring exceptional reliability for Grafana Cloud's highest-SLA customers. The role operates at the intersection of customer needs, production systems, and product engineering.

Grafana Labs is a 100% remote company. Candidates are sought from the UK, Sweden, Spain, or Germany.

Responsibilities

  • Partner closely with product engineering squads through an embedded model.
  • Own production reliability for high-SLA and complex customer environments.
  • Design and implement automation to scale reliability practices and eliminate toil.
  • Ensure customers meet SLO targets.
  • Define and evolve per-tenant SLOs and reliability models.
  • Proactively reduce SLO burn and prevent repeat incidents.
  • Serve as a primary escalation point and participate in on-call for relevant incidents.
  • Lead customer-impacting incident response and post-incident reviews.
  • Contribute to design documents and code reviews.
  • Influence feature design to ensure production scalability and operability.
  • Improve alert quality and reduce noisy escalations.
  • Investigate ways to reduce SLO budget burn through monitoring, automation, self-healing, and auto-scaling improvements.
  • Improve customer observability within their environments.
  • Design and implement reliable, scalable solutions for rapidly increasing demand.
  • Develop fault-tolerant design patterns across the service lifecycle.
  • Collaborate with engineering leaders on product strategy, roadmaps, and technical designs.
  • Teach Site Reliability Engineering practices and promote reliability best practices early in feature development.
  • Participate in incident response from investigation through resolution, post-incident review, and customer communication where necessary.
  • Use AI coding assistants for prototyping, test generation, refactoring, documentation, and incident follow-ups within security guidelines and with strong code review standards.

Requirements

  • 8+ years of engineering experience, including 4+ years in SRE, CRE, or production engineering. Formal customer reliability engineering experience is strongly preferred.
  • Strong Kubernetes experience in AWS, GCP, or Azure.
  • Familiarity with infrastructure-as-code tooling such as Helm, Terraform, and Jsonnet.
  • Strong technical leadership experience, including leading project teams, mentoring engineers, and acting as a force multiplier.
  • Experience operating multi-tenant systems in production.
  • Strong experience designing and implementing SLOs.
  • Experience with one or more programming languages, such as Go, Python, or Java.
  • Experience with Linux operating system internals and knowledge of networking, cloud storage, and scaling.
  • Excellent problem-solving and troubleshooting skills.
  • Experience participating calmly and actively in blame-free incident response, following up on actions, and writing high-quality PIRs or post-mortems.
  • Ability to reason about performance, scaling, and failure modes.
  • Ability to work autonomously and exercise self-direction within an engineering team.
  • Ability to partner deeply with product engineering teams.
  • Intellectual curiosity, transparency, a bias toward action, and a collaborative attitude.

Benefits

  • In Germany, the base compensation range is €109,709–€131,651 per year. Actual compensation may vary based on level, experience, and skill set.
  • Equity, bonus where applicable, and additional benefits.
  • 100% remote global culture.
  • In-person onboarding.
  • Global annual leave policy of 30 days per year, including 3 Grafana Shutdown Days, subject to local legislation.
  • Career growth pathways and an autonomy-focused work environment.

More jobs at Grafana Labs

Similar jobs