Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AWS @ 2
Azure @ 2
Distributed Systems @ 2
GCP @ 2
Go
Grafana
Helm @ 3
Kubernetes @ 2
Microservices
Node.js @ 3
Terraform @ 3
TypeScript @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Grafana Labs is seeking a Backend Engineer to join the Application Core Services (AppCore) group within the Platform Foundations department. AppCore develops systems that support customer workflows and internal business operations, including billing, provisioning, cloud marketplace integrations, and customer account management.
The AppCore Stacks squad owns the systems that create, configure, reconcile, migrate, and operate Grafana Cloud stacks at scale. The role focuses on control-plane services and workflows that keep stack state aligned across grafana.com, Stack State Service (SSS), Hosted Grafana, cloud regions, and Grafana Cloud infrastructure.
Responsibilities
- Design, build, and operate reconciliation systems, including the SSS backend, to track desired stack state and detect and repair drift across stack templates, grafana.com state, Hosted Grafana, and customer stack configuration.
- Collaborate across SSS, grafana.com, and deployment configurations to keep stack lifecycle workflows reliable, observable, and resilient.
- Reduce deployment complexity and contribute to the Stack Config Reconciliation project.
- Manage rollout mechanisms for provisioned plugins, dashboards, data sources, Grafana versions, release channels, and stack-level configuration.
- Support new region and cluster rollouts, including the operational paths required to bring stacks online safely in new Grafana Cloud regions.
- Improve incident response and recovery for stack misalignment, reconciliation failures, plugin rollout issues, and Hosted Grafana integration failures.
- Partner with Product, Hosted Grafana, Infrastructure, Support, and adjacent AppCore squads on customer-impacting stack lifecycle work.
- Contribute to roadmap planning, technical design, OnCall improvements, and long-term simplification of stack operations.
- Own the production behavior of the systems you build by improving runbooks, dashboards, alerts, reconciliation safety, rollout controls, and recovery procedures.
- Debug across service boundaries and make careful changes to systems that affect customer stacks.
- Write efficient, readable, maintainable, and well-tested code.
- Implement new microservices and systems.
- Collaborate across teams and departments to reach consensus on proposed solutions.
- Coordinate with Product and UX when needed.
- Respond to customer requests and feedback.
- Participate in a follow-the-sun OnCall rotation when ready.
- Participate in roadmap planning and prioritization decisions.
Requirements
- At least one year of fully remote work experience.
- Experience working on a SaaS platform and familiarity with distributed systems concepts such as scalability, multi-tenancy, and high availability.
- Professional experience with Golang and willingness to work across backend services and application code.
- Care for developer experience, user experience, and product quality.
- Experience contributing to projects from initial brainstorming through customer delivery.
- Ability to write clean, well-tested software that other engineers can understand, operate, and maintain.
- Ability to take on well-defined tasks, break them down, and execute iteratively while gathering feedback.
- Willingness to collaborate across teams and align work with other squads and external stakeholders.
- Familiarity with Kubernetes in AWS, GCP, or Azure.
- Exposure to infrastructure-as-code tooling such as Helm, Terraform, or Jsonnet.
- Experience participating in blameless incident response and contributing to post-incident reviews.
Bonus Points
- Experience with TypeScript or Node.js.
- Experience with Kubernetes control-plane patterns, operators, reconcilers, or desired-state systems.
- Experience with Jsonnet/Tanka, Terraform, Flux, Argo, or similar deployment and configuration tooling.
- Experience with SaaS provisioning, tenancy, regional expansion, plugin rollout, or customer lifecycle systems.
- Experience with incident response involving configuration drift, partial failure, or cross-service state mismatch.
Compensation And Rewards
- United Kingdom compensation range: GBP 72,000–GBP 90,000.
- Compensation may vary based on level, experience, and skill set.
- The role includes Restricted Stock Units (RSUs).
Benefits
- 100% remote, global culture.
- In-person onboarding.
- Global annual leave policy of 30 days per annum, including three Grafana Shutdown Days, subject to local legislation.
- Career growth pathways.
- Open-source culture and transparent communication.
The role is available to candidates located in the United Kingdom, Germany, Spain, Ireland, and Sweden.