Staff Software Engineer - Platform, Application Core Services
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Distributed Systems @ 6
Go @ 7
Grafana @ 4
Kubernetes @ 4
Node.js
Observability
Terraform @ 6
TypeScript
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture.
Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand, and act on all their disparate data to move at the speed of their ambitions.
This is a remote opportunity and we would be interested in applicants located in Canadian time zones (EST + CST only at this time).
Responsibilities
Application Core Services (AppCore) builds the control plane behind Grafana Cloud, designing the systems that create, configure, and operate thousands of customer environments safely and reliably.
Working at the intersection of product, platform, and operations, the team develops the services that power critical customer and internal workflows, from provisioning and cloud marketplace integrations to billing and customer lifecycle management. As Grafana Cloud continues to scale, we're solving complex distributed systems challenges around reliability, automation, and operational efficiency.
In this role you'll:
- Design and build distributed backend systems that keep customer stack state consistent across Grafana Cloud.
- Develop reconciliation and automation workflows that improve reliability, reduce operational complexity, and safely manage infrastructure changes.
- Lead technical initiatives spanning multiple services and partner closely with Product, Infrastructure, Hosted Grafana, and adjacent platform teams.
- Improve deployment workflows, regional expansion, rollout safety, and incident recovery across thousands of customer environments.
- Own the production systems you build, including observability, operational tooling, on-call, and continuous reliability improvements.
- Participate in a global follow-the-sun on-call rotation and invest heavily in AI-assisted development with access to modern coding tools and frontier models.
Requirements
You'll likely be successful if you:
- Have experience building and operating large-scale SaaS or cloud platforms.
- Have strong backend engineering experience in Go or another systems language, and are comfortable learning new technologies.
- Have experience with Kubernetes, cloud infrastructure, and infrastructure-as-code.
- Enjoy solving problems around distributed systems, eventual consistency, stateful services, and system reliability.
- Can independently lead projects, navigate ambiguity, and influence technical direction across teams.
- Care about writing clean, maintainable software and improving the developer experience.
- Have experience participating in production operations, incident response, and on-call.
- Are comfortable working in a remote-first, highly collaborative environment.
Bonus Points For:
- Kubernetes operators, reconcilers, or control plane development.
- Go and TypeScript/Node.js.
- Terraform, Flux, Argo, Jsonnet, Tanka, or similar infrastructure tooling.
- SaaS provisioning, customer lifecycle systems, or multi-tenant platforms.
- Problems related to configuration drift, partial failures, or cross-service consistency.
Benefits
- Benefits include equity, bonus (if applicable) and other benefits listed here: https://grafana.com/about/careers/#jobs.
- All roles include Restricted Stock Units (RSUs), giving every team member ownership in Grafana Labs' success.
In-Person onboarding and balance:
- In-Person onboarding.
- Global annual leave policy of 30 days per annum, with 3 days of annual leave entitlement reserved for Grafana Shutdown Days.
We will comply with local legislation where applicable.