Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
AWS @ 2
Azure @ 2
Distributed Systems @ 2
GCP @ 2
Go
Grafana
Helm @ 3
Kubernetes @ 2
Microservices
Node.js @ 3
Terraform @ 3
TypeScript
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Grafana Labs is seeking a Backend Engineer to join the Application Core Services team within the Platform Foundations department. The team builds and operates the systems that create, configure, reconcile, migrate, and operate Grafana Cloud stacks at scale.
The role focuses on control-plane services and workflows that keep stack state aligned across grafana.com, Stack State Service, Hosted Grafana, cloud regions, and Grafana Cloud infrastructure. You will work on systems involving stateful services, eventual consistency, reconciliation loops, scalability, reliability, and operational clarity.
Responsibilities
- Design, build, and operate reconciliation systems, including the Stack State Service backend.
- Track desired stack state, detect drift, and repair discrepancies across stack templates, grafana.com, Hosted Grafana, and customer stack configuration.
- Collaborate across Stack State Service, grafana.com, and deployment configurations to maintain reliable, observable, and resilient stack lifecycle workflows.
- Reduce deployment complexity and contribute to the Stack Config Reconciliation project.
- Manage rollout mechanisms for plugins, dashboards, data sources, Grafana versions, release channels, and stack-level configuration.
- Support new region and cluster rollouts and the operational processes required to bring stacks online safely.
- Improve incident response and recovery for stack misalignment, reconciliation failures, plugin rollout issues, and Hosted Grafana integration failures.
- Partner with Product, Hosted Grafana, Infrastructure, Support, and adjacent Application Core Services squads.
- Contribute to roadmap planning, technical design, OnCall improvements, runbooks, dashboards, alerts, rollout controls, and recovery procedures.
- Write efficient, readable, maintainable, and well-tested code.
- Implement new microservices and systems.
- Collaborate across teams and departments to reach consensus on technical solutions.
- Participate in product and UX collaboration, roadmap planning, prioritization, customer feedback, and a follow-the-sun OnCall rotation when ready.
- Use AI-assisted and agentic development practices across design, implementation, testing, documentation, and operations where appropriate.
Requirements
- At least one year of fully remote work experience.
- Experience working on a SaaS platform and familiarity with distributed systems concepts such as scalability, multi-tenancy, and high availability.
- Professional experience with Golang and willingness to work across backend services and application code.
- Experience contributing to projects from initial brainstorming through delivery to customers.
- Ability to break down well-defined tasks and execute iteratively while gathering feedback.
- Strong focus on developer experience, user experience, and product quality.
- Ability to collaborate across teams and align work with other squads and external stakeholders.
- Familiarity with Kubernetes in AWS, GCP, or Azure.
- Exposure to infrastructure-as-code tooling such as Helm, Terraform, or Jsonnet.
- Experience participating in blameless incident response and post-incident reviews.
Bonus Skills
- TypeScript or Node.js experience.
- Experience with Kubernetes control-plane patterns, operators, reconcilers, or desired-state systems.
- Experience with Jsonnet/Tanka, Terraform, Flux, Argo, or similar deployment and configuration tooling.
- Experience with SaaS provisioning, tenancy, regional expansion, plugin rollout, or customer lifecycle systems.
- Experience responding to configuration drift, partial failure, or cross-service state mismatch.
Compensation
The base compensation range in Spain is EUR 65,000–EUR 81,000 per year. Actual compensation may vary based on level, experience, and skillset. The role also includes Restricted Stock Units (RSUs).
Benefits
- 100% remote work in a global culture.
- In-person onboarding.
- Global annual leave policy of 30 days per year, including three Grafana Shutdown Days.
- Career growth opportunities.
- Open-source, transparent, and empowerment-focused working culture.
The role is available to candidates located in the UK, Germany, Spain, Ireland, and Sweden. Grafana Labs is an equal opportunities employer.