Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
AWS @ 4
Azure @ 4
DevOps @ 4
Docker @ 4
GCP @ 4
GenAI
Generative AI @ 4
Grafana @ 4
Kubernetes @ 4
LLM
Observability @ 4
Prompt Engineering @ 4
Security @ 7
Terraform @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Grafana Labs is seeking a Senior AI Engineer to build AI-driven features that help users understand, respond to, and improve their systems through observability data. This remote opportunity is currently open to applicants in Canada time zones only.
The role focuses on developing, testing, and shipping AI-powered features that automate infrastructure and observability workflows, reduce toil, and expand the capabilities of AI agents for incident response. The successful candidate will have a strong software engineering background, a practical approach to AI, and a willingness to experiment, iterate quickly, and scale impactful solutions.
Responsibilities
- Develop and deliver high-performance AI features that help users detect, triage, and resolve incidents using observability data and tools.
- Rapidly prototype, test, validate, ship, and evolve LLM- and agent-powered workflows for incident lifecycle management and automated analysis.
- Collaborate with data analysts, product managers, and designers to shape AI-driven product features.
- Integrate agentic components with internal tools, alerting systems, runbooks, and developer workflows.
- Use AI and automation tools to improve product functionality and development workflows.
- Communicate effectively and contribute across teams in a dynamic, collaborative environment.
- Take ownership of AI solutions, ensuring they are innovative, scalable, maintainable, and aligned with real user workflows.
- Use modern AI coding assistants, including frontier models from OpenAI, Anthropic, and Google, within security guidelines and supported by strong code review and quality standards.
Requirements
- Strong experience building production software systems, including backend and/or full-stack systems.
- Familiarity with AI technologies and frameworks, with a focus on practical, high-quality solutions.
- Experience with LLMs, prompt engineering, and building applications powered by generative AI.
- A proven track record of delivering software that reached production and is actively used by customers or users.
- Experience working in cloud-native environments such as AWS, GCP, or Azure.
- Experience using observability tools to understand and troubleshoot system behavior.
- Ability to work independently, handle ambiguity, define scope, communicate effectively, collaborate cross-functionally, and drive projects forward.
Bonus Qualifications
- Experience building or working with agent frameworks or multi-agent workflows.
- Experience with infrastructure or DevOps tooling such as Kubernetes, Docker, Terraform, or similar deployment technologies.
- Familiarity with model fine-tuning techniques.
- Experience building observability tooling.
Benefits
- 100% remote work and a global collaborative culture.
- Restricted Stock Units (RSUs).
- Company-funded usage budget for AI coding assistants.
- Career growth pathways.
- In-person onboarding.
- Global annual leave policy of 30 days per year, including three Grafana Shutdown Days, subject to local legislation.
More jobs at Grafana Labs
Software Engineer - Platform Metal | Ireland | Remote
Grafana Labs · Spain, United Kingdom, Ireland
EUR 81,600-97,900 per year
Software Engineer - Platform Metal | Spain | Remote
Grafana Labs · Spain, United Kingdom, Ireland
EUR 65,300-78,300 per year
Software Engineer - Platform Metal | UK | Remote
Grafana Labs · Spain, United Kingdom, Ireland
GBP 72,200-86,600 per year
Senior Solutions Engineer
Grafana Labs · Tokyo, Japan
JPY 14,000,000-18,500,000 per year
Senior Solutions Engineer | West Coast | Remote
Grafana Labs · United States
USD 204,000-254,000 per year
Similar jobs
Staff AI Engineer - Grafana AI/ML
Grafana Labs · Canada
CAD 186,400-230,000 per year
Senior AI Engineer - Grafana AI/ML
Grafana Labs · United States
USD 127,700-203,900 per year
Staff AI Engineer - Grafana AI/ML
Grafana Labs · United States
USD 175,000-220,000 per year
Principal ML Solutions Architect - Token Factory
Nebius · United States
USD 208,000-261,000 per year
ML Solution Architect (Early Talent)
Nebius · United States
USD 102-126 per hour
Senior ML Solutions Architect - Token Factory
Nebius · United States
USD 210,000-260,000 per year
Staff Backend Engineer - Second Horizon | Canada | Remote
Grafana Labs · Canada
CAD 186,400-223,600 per year
Staff Backend Engineer - Second Horizon | US | Remote
Grafana Labs · United States, Canada
USD 175,000-210,000 per year