Staff Software Engineer - Adaptive Telemetry, Databases

📍 Canada
CAD 186,400-223,600 per year
SENIOR
✅ Remote

Tech Stack

AI @ 4 Communication @ 6 Compliance Debugging Distributed Systems @ 6 Grafana @ 3 IaC @ 4 Kafka @ 3 Kubernetes @ 4 Leadership @ 4 Microservices @ 4 Observability @ 3 Prometheus @ 3 Python @ 6 Rust @ 6 Technical Leadership @ 4

Details

Responsibilities

  • Drive technical strategy and roadmap. Proactively define the architectural vision, prioritize work that unlocks major product or platform improvements, and influence product and engineering decisions.
  • Lead end-to-end delivery of large, cross-functional projects. Own planning, design, execution, rollout and long-term operation of large initiatives.
  • Own architecture, reliability, performance and cost for critical systems. Make pragmatic architecture choices that balance scalability, availability, latency and cost while ensuring systems remain maintainable and evolvable.
  • Define SLOs/SLIs and lead incident response. Establish measurable reliability targets, run high-severity incident response, lead blameless post-mortems, and drive systemic fixes and automation to prevent recurrence.
  • Improve observability, automation and operational readiness. Champion telemetry, alerting, runbooks, capacity planning and automation efforts that reduce toil, speed debugging and lower MTTR.
  • Align stakeholders and remove blockers. Coordinate across Product, Design and other teams to align priorities, negotiate tradeoffs, and unblock delivery for large initiatives.
  • Mentor and grow engineering talent. Coach senior and mid-level engineers, lead design reviews, raise engineering standards, and help teammates make sound technical tradeoffs.
  • Represent engineering internally and externally. Communicate technical strategy clearly to non-engineering stakeholders and represent the team in cross-team planning.

Requirements

  • Proven delivery of large distributed systems. Experience shipping and operating complex systems that span multiple teams, with clear evidence of technical leadership and impact.
  • Strong systems-design instincts. Deep understanding of tradeoffs around latency, consistency, availability, scaling and cost.
  • Hands-on cloud and platform experience. Solid experience with cloud-native architectures (microservices, containers/Kubernetes, IaC) and the operational practices that keep them healthy.
  • Reliability and performance ownership. Comfortable defining SLOs/SLIs, doing capacity planning, tuning performance, and driving reliability work end-to-end.
  • Excellent coding and design skills. Write clear, maintainable, well-tested code and lead technical designs — Go is used, but Python/C/C++/Rust or similar translate well.
  • Comfort with AI-assisted development. Curious and comfortable using AI-powered developer tools and ideally practical experience folding them into a team’s workflow.
  • Experience with messaging and telemetry. Familiarity with streaming/messaging systems (e.g., Kafka) and observability tooling (Prometheus/Grafana or equivalents).
  • Influence without authority. Ability to align cross-functional stakeholders, set priorities and drive outcomes in a remote-first environment.
  • Strong communicator. Clear written and verbal communication that works across engineers and non-technical stakeholders.

Benefits

  • Equity

  • Bonus (if applicable)

  • Other benefits listed here: https://grafana.com/about/careers/#jobs

  • 100% Remote, Global Culture

  • Transparent Communication

  • Innovation-Driven

  • Open Source Roots

  • Empowered Teams

  • Career Growth Pathways

  • Approachable Leadership

  • In-Person onboarding (for day 1 onboarding)

  • Balance is Key: global annual leave policy of 30 days per annum, with 3 reserved for Grafana Shutdown Days (compliance with local legislation where applicable)

More jobs at Grafana Labs

Similar jobs