Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
AWS @ 7
Azure @ 7
GCP @ 7
Go @ 4
Grafana
Helm @ 3
Java @ 4
Kubernetes @ 7
Leadership @ 7
Linux @ 4
Mentoring @ 7
Networking @ 4
Observability
Python @ 4
SRE @ 9
Security
Technical Leadership @ 7
Terraform @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Grafana Labs is looking for a Staff Software Engineer - SRE to support high-value Grafana Cloud customers by increasing the reliability of cloud databases based on Mimir, Loki, Tempo, and Pyroscope. These SaaS databases run across AWS, GCP, and Azure in all regions.
The SRE team is embedded within the Mimir, Loki, and Tempo squads and focuses on delivering exceptional reliability for high-SLA customers. The role operates at the intersection of customer needs, production systems, and product engineering.
Grafana Labs is a fully remote company. Candidates are sought from the UK, Sweden, Spain, or Germany.
Responsibilities
- Partner closely with product engineering squads through an embedded team model.
- Own production reliability for high-SLA and complex customer environments.
- Design and implement automation to scale reliability practices and eliminate toil.
- Ensure customers meet SLO targets and define and evolve per-tenant SLOs and reliability models.
- Proactively reduce SLO burn and prevent repeat incidents.
- Serve as a primary escalation point and participate in on-call rotations.
- Lead customer-impacting incident response and post-incident reviews.
- Contribute to design documents, code reviews, and technical designs.
- Influence feature design to ensure production scalability and operability.
- Improve alert quality and reduce noisy escalations.
- Investigate ways to reduce SLO budget burn through monitoring, automation, self-healing, and autoscaling improvements.
- Improve customer observability within their environments.
- Design and implement fault-tolerant solutions that support reliability and scalability throughout the service lifecycle.
- Collaborate with engineering leaders to define product strategy, roadmaps, and technical designs.
- Teach Site Reliability Engineering practices and communicate best practices for new features and functionality.
- Participate in incident response from investigation through resolution, post-incident review, and customer communication via bridge calls when necessary.
Requirements
- 8+ years of engineering experience, including 4+ years in SRE, CRE, or production engineering. Formal customer reliability engineering experience is strongly preferred.
- Strong Kubernetes experience in AWS, GCP, or Azure.
- Familiarity with infrastructure-as-code tooling such as Helm, Terraform, and Jsonnet.
- Strong technical leadership experience, including leading teams through projects, mentoring engineers, and acting as a force multiplier.
- Experience operating multi-tenant systems in production.
- Strong experience designing and implementing SLOs.
- Experience with one or more programming languages, such as Go, Python, or Java.
- Experience with Linux operating system internals and knowledge of networking, cloud storage, and scaling.
- Excellent problem-solving and troubleshooting skills.
- Experience participating calmly and actively in blame-free incident response, following up on actions, and writing high-quality post-incident reviews or post-mortems.
- Ability to reason about performance, scaling, and failure modes.
- Ability to work autonomously and collaborate deeply with product engineering teams.
- Intellectual curiosity, transparency, a bias toward action, and a kind, collaborative approach.
Benefits
- Base compensation in Sweden: SEK 878,578–SEK 1,054,294 per year.
- Equity, bonus where applicable, and other benefits.
- 100% remote global culture.
- Company-funded usage budget for modern AI coding assistants, within security guidelines.
- In-person onboarding.
- 30 days of annual leave per year, including 3 Grafana Shutdown Days where applicable.
- Career growth pathways and an innovation-driven, high-trust work environment.
More jobs at Grafana Labs
Senior Solutions Engineer
Grafana Labs · Tokyo, Japan
JPY 14,000,000-18,500,000 per year
Senior Solutions Engineer | West Coast | Remote
Grafana Labs · United States
USD 204,000-254,000 per year
Associate Observability Architect | PST | Remote
Grafana Labs · United States
USD 139,000-167,000 per year
Senior Platform Engineer - Platform Metal
Grafana Labs · Spain, United Kingdom, Ireland
GBP 91,800-110,100 per year
Senior Solutions Engineer | France | Arabic Speaker | Remote
Grafana Labs · France
EUR 115,800-138,000 per year
Similar jobs
Staff Software Engineer - Databases SRE | Ireland | Remote
Grafana Labs · Ireland, Spain, Sweden, Germany, United Kingdom
EUR 117,600-141,100 per year
Staff Software Engineer - Databases SRE
Grafana Labs · Germany, Sweden, Spain, United Kingdom
EUR 109,700-131,700 per year
Staff Software Engineer - Databases SRE
Grafana Labs · Spain, Sweden, Germany, United Kingdom
EUR 94,000-112,800 per year
Staff Software Engineer - Databases SRE
Grafana Labs · United Kingdom, Sweden, Germany, Spain
GBP 104,000-124,800 per year
Senior Platform Engineer
Collibra · United States
USD 168,000-210,000 per year
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Service Reliability Engineer
Nvidia · United States
USD 168,000-333,500 per year
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year