Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
AWS @ 7
Azure @ 7
GCP @ 7
Go @ 4
Grafana @ 9
Helm @ 3
Java @ 4
Kubernetes @ 7
Leadership @ 7
Linux @ 4
Mentoring @ 7
Networking @ 4
Observability
Python @ 4
SRE @ 9
Security
Technical Leadership @ 7
Terraform @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Grafana Labs is seeking a Staff Software Engineer - SRE to support high-value Grafana Cloud customers by increasing the reliability of cloud databases based on Mimir, Loki, Tempo, and Pyroscope. These databases are provided as a SaaS product across AWS, GCP, and Azure in all regions.
The SRE team is embedded within the Mimir, Loki, and Tempo squads and is responsible for ensuring exceptional reliability for Grafana Cloud's highest-SLA customers. This is a remote opportunity for candidates based in the United Kingdom, Sweden, Spain, or Germany.
Responsibilities
- Partner closely with product engineering squads using an embedded model.
- Own production reliability for high-SLA and complex customer environments.
- Design and implement automation to scale reliability practices and eliminate toil.
- Define and evolve per-tenant SLOs and reliability models.
- Ensure customers meet SLO targets and proactively reduce SLO burn.
- Serve as a primary escalation point and participate in on-call rotations.
- Lead customer-impacting incident response and post-incident reviews.
- Contribute to design documents and code reviews.
- Influence feature design to ensure production scalability and operability.
- Improve alert quality and reduce noisy escalations.
- Improve customer observability within their environments.
- Design and implement fault-tolerant, reliable, and scalable solutions.
- Collaborate with engineering leaders to define product strategy, roadmaps, and technical designs.
- Teach Site Reliability Engineering practices and promote reliability early in the development lifecycle.
- Participate in incident investigation, resolution, post-incident reviews, and customer communication when required.
Grafana Labs invests in developer productivity and supports the use of modern AI coding assistants within security guidelines, including company-funded usage and access to frontier models. AI-assisted development is paired with code review and quality standards.
Requirements
- 8+ years of engineering experience, including 4+ years in SRE, CRE, or production engineering; formal customer reliability engineering experience is strongly preferred.
- Strong Kubernetes experience on AWS, GCP, or Azure.
- Familiarity with infrastructure-as-code tools such as Helm, Terraform, and Jsonnet.
- Strong technical leadership experience, including leading teams through projects, mentoring engineers, and acting as a force multiplier.
- Experience operating multi-tenant production systems.
- Strong experience designing and implementing SLOs.
- Experience with one or more programming languages, such as Go, Python, or Java.
- Experience with Linux operating system internals and knowledge of networking, cloud storage, and scaling.
- Excellent problem-solving and troubleshooting skills.
- Experience participating in blame-free incident response, following up on actions, and writing high-quality post-incident reviews or post-mortems.
- Ability to reason about performance, scaling, and failure modes.
- Ability to work autonomously and self-direct within an engineering team.
- Ability to partner deeply with product engineering teams.
- Intellectual curiosity, transparency, bias toward action, and a collaborative approach.
Benefits
- Equity, bonus where applicable, and other benefits.
- 100% remote work and a global culture.
- Career growth pathways.
- In-person onboarding.
- 30 days of annual leave per year, including three Grafana Shutdown Days, subject to local legislation.
Compensation
The UK base compensation range is £103,958–£124,750 per year. Actual compensation may vary based on level, experience, and skill set. Compensation ranges and benefits are country-specific.