Staff Backend Engineer - Grafana Enterprise | US | Remote
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API @ 4
CI/CD @ 4
Communication @ 7
Distributed Systems @ 7
Go
Grafana @ 4
IaaS @ 4
Kubernetes @ 4
MySQL @ 7
Observability @ 4
OpenTelemetry @ 4
PostgreSQL @ 7
Prometheus @ 4
React @ 4
SRE
Security
TypeScript @ 4
gRPC @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Grafana Labs is seeking an experienced software engineer to design and build the backend systems powering Grafana Enterprise, a composable observability platform used by large-scale operators and cloud service providers. The role focuses on distributed systems, reliable and scalable backend infrastructure, enterprise features, security, robustness, flexibility, multitenancy, and interoperability.
The backend is primarily built in Go. Engineers contribute to open-source communities and work with software engineers, site reliability engineers, platform operators, product teams, UX teams, frontend engineers, and enterprise customers. Grafana Labs is a remote-only company, with meetings generally occurring between 14:00 and 17:00 UTC.
Responsibilities
- Design and lead the development of backend services, distributed systems, and enterprise features at scale.
- Architect and implement distributed backend services in Go, focusing on correctness, observability, and performance.
- Design APIs and service contracts used by enterprise operators and cloud service providers.
- Drive projects from ideation through development, production deployment, and operations.
- Improve the scalability, reliability, security, and multitenancy of the Grafana platform.
- Collaborate with Product, UX, and frontend engineering teams to deliver complete solutions.
- Engage directly with large enterprise customers and cloud service providers to understand requirements and translate them into robust engineering solutions.
- Own the operational health of the platform through weekday 12-hour-by-5-day and weekend 24-hour-by-2-day on-call rotations.
- Drive continuous improvement of engineering and operational practices.
- Hire, develop, and mentor engineers.
- Advocate for customer needs throughout the development lifecycle.
- Participate in a communicative, remote-first engineering culture using written communication and video calls.
- Use AI coding assistants for prototyping, test generation, refactoring, documentation, and incident follow-ups within security and code-quality guidelines.
Requirements
- Deep professional experience writing production services from ideation through production operations at scale.
- Strong distributed systems fundamentals, including replication, consistency models, partitioning, fault tolerance, and associated scaling trade-offs.
- Experience designing and operating large-scale, high-traffic, high-availability, or multi-tenant systems, ideally for infrastructure, observability, or software delivery platforms.
- Professional experience building and consuming gRPC and Protocol Buffers APIs and designing service contracts across service boundaries.
- Strong database skills with PostgreSQL and/or MySQL, including schema design, query optimization, and schema migrations at scale.
- Experience with large-scale CI/CD systems and build tooling, including designing, operating, or integrating with continuous delivery pipelines.
- Experience with Kubernetes and containerized deployment environments, including stateful workloads and multi-tenant clusters.
- Experience with observability tooling such as OpenTelemetry, Prometheus metrics, structured logging, and distributed tracing.
- Familiarity with dependency injection patterns such as Google Wire and clean, testable service architecture.
- Strong teamwork, written communication, interpersonal skills, customer focus, and ability to solve complex distributed systems problems.
Nice-to-have Technical Competencies
- TypeScript and React experience.
- Experience with Grafana's LGTM+ observability stack: Loki, Mimir, Tempo, Pyroscope, and Alloy.
- Experience at or building for large-scale cloud service providers, IaaS providers, or global enterprises with demanding SLA requirements.
- Experience designing or operating large-scale build infrastructure, artifact registries, distributed build caches, hermetic build systems such as Bazel, or developer platform tooling.
Compensation
The United States compensation range is $174,986–$209,983 USD per year. Actual compensation may vary based on level, experience, and skillset. The role includes Restricted Stock Units (RSUs).
Benefits
- 100% remote, global work culture.
- In-person onboarding.
- Global annual leave policy of 30 days per year, including three Grafana Shutdown Days where applicable.
- Career growth pathways and an innovation-driven, high-trust environment.
- Open-source culture and transparent communication.