Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 7
API @ 4
CI/CD @ 4
Distributed Systems @ 7
Go
Grafana @ 4
IaaS @ 4
Kubernetes @ 4
MySQL @ 7
Observability @ 4
OpenTelemetry @ 4
PostgreSQL @ 7
Prometheus @ 4
React @ 4
Security @ 4
TypeScript @ 4
gRPC @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Grafana Labs is looking for an experienced Staff Backend Engineer to design and build the backend systems powering Grafana Enterprise, a composable observability platform for large-scale operators with security and regulatory requirements. The role focuses on distributed systems, reliable and scalable backend infrastructure, enterprise features, and open-source observability software. The position is fully remote and open to candidates in Canada and the United States.
The backend is built with Go, and the role involves working on tools used by operators of large observability and software delivery stacks. Engineers collaborate across Grafana Labs and directly with major cloud service providers and global enterprises to solve demanding distributed systems challenges.
Responsibilities
- Design and lead the development of backend services, distributed systems, and enterprise features at scale.
- Architect and implement distributed backend services in Go, focusing on correctness, observability, and performance.
- Design APIs and service contracts used by enterprise operators and cloud service providers.
- Drive projects from initial ideation through development and production operations.
- Improve the scalability, reliability, security, and multi-tenancy of the Grafana platform.
- Collaborate with Product, UX, and frontend engineers to deliver complete end-to-end solutions.
- Engage directly with enterprise customers and cloud service providers to understand requirements and translate them into robust engineering solutions.
- Contribute to open-source communities and advocate for customers throughout the development lifecycle.
- Drive continuous improvement of engineering and operational practices.
- Participate in weekday 12-hour by 5-day and separate weekend 24-hour by 2-day on-call rotations.
- Hire and develop engineers and contribute to the future of observability.
- Use AI coding assistants for prototyping, test generation, refactoring, documentation, and incident follow-ups within security guidelines and with strong code review standards.
Requirements
- Deep professional experience writing production services from ideation through large-scale production operations.
- Strong distributed systems fundamentals, including replication, consistency models, partitioning, fault tolerance, and scalability trade-offs.
- Experience designing and operating large-scale, high-traffic, high-availability, or multi-tenant systems, ideally for infrastructure, observability, or software delivery platforms.
- Professional experience building and consuming gRPC and Protocol Buffers APIs and designing service contracts across service boundaries.
- Strong database skills with PostgreSQL and/or MySQL, including schema design, query optimization, and schema migrations at scale.
- Experience with large-scale CI/CD systems and build tooling.
- Experience designing, operating, or integrating continuous delivery pipelines for large engineering organizations or external operators.
- Experience with Kubernetes and containerized deployment environments, including stateful workloads and multi-tenant clusters.
- Experience with observability tooling, including OpenTelemetry, Prometheus metrics, structured logging, and distributed tracing.
- Familiarity with dependency injection patterns such as Google Wire and clean, testable service architecture.
- Strong written, interpersonal, and remote collaboration skills.
- Ability to work effectively with engineering teams, understand customer needs, and break down complex distributed systems problems.
Nice-to-have Technical Competencies
- Experience with TypeScript and React.
- Experience with Grafana's LGTM+ observability stack: Loki, Mimir, Tempo, Pyroscope, and Alloy.
- Experience at or building for large-scale cloud service providers, IaaS providers, or global enterprises with demanding SLA requirements.
- Experience designing or operating large-scale build infrastructure, artifact registries, distributed build caches, hermetic build systems such as Bazel, or developer platform tooling.
Benefits
- Restricted Stock Units (RSUs) are included in all roles.
- Company-funded usage budget for modern AI coding assistants.
- 100% remote global culture.
- Career growth pathways.
- In-person onboarding.
- Global annual leave policy of 30 days per annum, including three Grafana Shutdown Days, subject to local legislation.
Compensation
The compensation range in Canada is $186,368–$223,642 CAD per year. Actual compensation may vary based on level, experience, and skill set.