Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AWS @ 7
Azure @ 7
GCP @ 7
Go @ 4
Grafana @ 1
Helm @ 3
Java @ 4
Kubernetes @ 7
Linux @ 4
Networking @ 4
Python @ 4
SRE @ 7
Terraform @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Grafana Labs is looking for a Staff Software Engineer - SRE to help support Grafana Cloud’s highest value customers by increasing the reliability of their Cloud databases based on Mimir, Loki, Tempo, and Pyroscope.
This is a remote opportunity (open to candidates from the UK, Sweden, Spain, or Germany).
Responsibilities
- Partner closely with product engineering squads (embedded model)
- Own production reliability for high-SLA and complex customer environments
- Design and implement automation to scale reliability practices
- Ensure customers meet their SLO targets
- Define and evolve per-tenant SLOs and reliability models
- Proactively reduce SLO burn to prevent repeat incidents
- Serve as a primary escalation point and on-call for relevant incidents
- Lead customer-impacting incident response and post-incident reviews
- Contribute to design docs and code reviews
- Influence feature design to ensure production scalability and operability
- Build automation to eliminate toil where needed
- Improve alert quality and reduce noisy escalations
Requirements
- 8+ years engineering experience, 4+ in SRE/CRE/production engineering
- Strong preference for candidates with formal customer reliability engineering experience
- Strong Kubernetes experience in AWS, GCP, or Azure
- Familiarity with infrastructure-as-code tooling (Helm, Terraform, Jsonnet, etc.)
- Experience operating multi-tenant systems in production
- Strong experience designing and implementing SLOs
- Experience with one or more programming languages (e.g., Go, Python, Java)
- Experience with Linux operating systems internals and some knowledge of networking, cloud storage, and scaling
- Excellent problem-solving and troubleshooting skills
- Experience participating in blame-free Incident Response, following up on actions, and writing high-quality PIRs (Post Incident Reviews)
- Ability to reason about performance, scaling, and failure modes
- Comfortable working with autonomy and self-direction within an engineering team
- Ability to partner deeply with product engineering teams
- Intellectual curiosity, defaulting to transparency, high bias towards action, and kindness
Benefits
- Equity and bonus (if applicable), plus other benefits listed here
- 100% remote, global culture
- Open-source roots and transparent communication
- In-person onboarding
- Global annual leave policy of 30 days per annum, including 3 Grafana Shutdown Days reserved from annual leave
Compensation
In Germany, the base compensation range is €109,709 - €131,651 (country-specific ranges apply).
More jobs at Grafana Labs
Senior Software Engineer - Databases, SRE
Grafana Labs · United States
USD 154,400-185,300 per year
Senior Software Engineer - Databases, SRE
Grafana Labs · Canada
CAD 164,500-197,400 per year
Staff Product Designer, Incident Response Management
Grafana Labs · United States, Canada
USD 162,000-202,000 per year
Staff Product Designer, Incident Response Management (IRM)
Grafana Labs · Canada
CAD 164,000-205,000 per year
Staff Ai Engineer - 2nd Horizon
Grafana Labs · Spain, Ireland, Sweden, Germany, United Kingdom
SEK 878,000-1,100,000 per year
Similar jobs
Staff Software Engineer - Databases SRE
Grafana Labs · Ireland
EUR 117,600-141,100 per year
Staff Software Engineer - Databases SRE
Grafana Labs · Sweden
SEK 878,600-1,054,300 per year
Staff Software Engineer - Databases SRE
Grafana Labs · Spain
EUR 94,000-112,800 per year
Staff Software Engineer - Databases SRE
Grafana Labs · United Kingdom
GBP 104,000-124,800 per year
Senior Platform Engineer
Collibra · United States
USD 168,000-210,000 per year
Staff Infrastructure Engineer
SentinelOne · United States
USD 132,000-215,000 per year
Senior Software Engineer - Public Cloud Engineering
Bloomberg · New York City, United States
USD 160,000-240,000 per year
Senior Systems Engineer, Storage - DGX Cloud
Nvidia · United States
USD 208,000-414,000 per year