Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 7
AWS @ 7
Azure @ 7
Design Patterns
GCP @ 7
Go @ 4
Grafana @ 9
Helm @ 3
Java @ 4
Kubernetes @ 7
Leadership @ 7
Linux @ 4
Mentoring @ 7
Networking @ 4
Observability
Python @ 4
SRE @ 9
Security @ 7
Technical Leadership @ 7
Terraform @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Grafana Labs is seeking a Staff Software Engineer - SRE to support its highest-value Grafana Cloud customers by improving the reliability of cloud databases based on Mimir, Loki, Tempo, and Pyroscope. These databases are provided as a SaaS product across AWS, GCP, and Azure regions.
The SRE team is embedded within the Mimir, Loki, and Tempo squads and is responsible for ensuring exceptional reliability for Grafana Cloud's highest-SLA customers. The role operates at the intersection of customer needs, production systems, and product engineering.
Grafana Labs is a 100% remote company. Candidates are sought from the UK, Sweden, Spain, or Germany.
Responsibilities
- Partner closely with product engineering squads through an embedded model.
- Own production reliability for high-SLA and complex customer environments.
- Design and implement automation to scale reliability practices and eliminate toil.
- Ensure customers meet SLO targets.
- Define and evolve per-tenant SLOs and reliability models.
- Proactively reduce SLO burn and prevent repeat incidents.
- Serve as a primary escalation point and participate in on-call for relevant incidents.
- Lead customer-impacting incident response and post-incident reviews.
- Contribute to design documents and code reviews.
- Influence feature design to ensure production scalability and operability.
- Improve alert quality and reduce noisy escalations.
- Investigate ways to reduce SLO budget burn through monitoring, automation, self-healing, and auto-scaling improvements.
- Improve customer observability within their environments.
- Design and implement reliable, scalable solutions for rapidly increasing demand.
- Develop fault-tolerant design patterns across the service lifecycle.
- Collaborate with engineering leaders on product strategy, roadmaps, and technical designs.
- Teach Site Reliability Engineering practices and promote reliability best practices early in feature development.
- Participate in incident response from investigation through resolution, post-incident review, and customer communication where necessary.
- Use AI coding assistants for prototyping, test generation, refactoring, documentation, and incident follow-ups within security guidelines and with strong code review standards.
Requirements
- 8+ years of engineering experience, including 4+ years in SRE, CRE, or production engineering. Formal customer reliability engineering experience is strongly preferred.
- Strong Kubernetes experience in AWS, GCP, or Azure.
- Familiarity with infrastructure-as-code tooling such as Helm, Terraform, and Jsonnet.
- Strong technical leadership experience, including leading project teams, mentoring engineers, and acting as a force multiplier.
- Experience operating multi-tenant systems in production.
- Strong experience designing and implementing SLOs.
- Experience with one or more programming languages, such as Go, Python, or Java.
- Experience with Linux operating system internals and knowledge of networking, cloud storage, and scaling.
- Excellent problem-solving and troubleshooting skills.
- Experience participating calmly and actively in blame-free incident response, following up on actions, and writing high-quality PIRs or post-mortems.
- Ability to reason about performance, scaling, and failure modes.
- Ability to work autonomously and exercise self-direction within an engineering team.
- Ability to partner deeply with product engineering teams.
- Intellectual curiosity, transparency, a bias toward action, and a collaborative attitude.
Benefits
- In Germany, the base compensation range is €109,709–€131,651 per year. Actual compensation may vary based on level, experience, and skill set.
- Equity, bonus where applicable, and additional benefits.
- 100% remote global culture.
- In-person onboarding.
- Global annual leave policy of 30 days per year, including 3 Grafana Shutdown Days, subject to local legislation.
- Career growth pathways and an autonomy-focused work environment.
More jobs at Grafana Labs
Senior Solutions Engineer
Grafana Labs · Tokyo, Japan
JPY 14,000,000-18,500,000 per year
Senior Solutions Engineer | West Coast | Remote
Grafana Labs · United States
USD 204,000-254,000 per year
Associate Observability Architect | PST | Remote
Grafana Labs · United States
USD 139,000-167,000 per year
Senior Platform Engineer - Platform Metal
Grafana Labs · Spain, United Kingdom, Ireland
GBP 91,800-110,100 per year
Senior Solutions Engineer | France | Arabic Speaker | Remote
Grafana Labs · France
EUR 115,800-138,000 per year
Similar jobs
Staff Software Engineer - Databases SRE | Ireland | Remote
Grafana Labs · Ireland, Spain, Sweden, Germany, United Kingdom
EUR 117,600-141,100 per year
Staff Software Engineer - Databases SRE
Grafana Labs · Sweden, Germany, Spain, United Kingdom
SEK 878,600-1,054,300 per year
Staff Software Engineer - Databases SRE
Grafana Labs · Spain, Sweden, Germany, United Kingdom
EUR 94,000-112,800 per year
Staff Software Engineer - Databases SRE
Grafana Labs · United Kingdom, Sweden, Germany, Spain
GBP 104,000-124,800 per year
Senior Platform Engineer
Collibra · United States
USD 168,000-210,000 per year
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Service Reliability Engineer
Nvidia · United States
USD 168,000-333,500 per year
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year