Staff AI Engineer - Grafana AI/ML
Build and ship AI-powered features for incident detection, triage, resolution, and observability workflows. The role requires strong production software engineering experience, practical LLM and GenAI expertise, cloud-native exposure, and a pragmatic approach to rapid experimentation.
About the job
Responsibilities
- Build and deliver high-performance AI features that help users detect, triage, and resolve incidents using observability data and tools.
- Prototype, test, ship, and evolve LLM- and agent-powered workflows for incident lifecycle management and automated analysis.
- Collaborate with data analysts, product managers, and designers on AI-driven product features.
- Integrate agentic components with internal tools, alerting systems, runbooks, and developer workflows.
- Use AI and automation tools to improve product functionality and development workflows.
- Own AI solutions end to end, ensuring they are innovative, scalable, maintainable, and aligned with user workflows.
Requirements
- Strong experience building production software systems, particularly backend and/or full-stack systems.
- Experience with LLMs, prompt engineering, and GenAI-powered applications.
- Proven delivery of production software actively used by customers.
- Exposure to cloud-native environments such as AWS, GCP, or Azure.
- Experience using observability tools to understand and troubleshoot system behavior.
- Ability to iterate quickly, work through ambiguity, take initiative, and communicate effectively across teams.
Nice-to-haves
- Experience with agent frameworks or multi-agent workflows.
- Experience with infrastructure or DevOps tooling such as Kubernetes, Docker, and Terraform.
- Familiarity with model fine-tuning techniques.
- Experience building observability tooling.
Compensation and Benefits
- Base compensation in Canada: CAD 186,368–CAD 230,000.
- Restricted Stock Units (RSUs) are included with all roles.
- Company-funded usage budget for modern AI coding assistants.
- Access to frontier models from OpenAI, Anthropic, and Google.
Skills
LLMs, Prompt Engineering, Generative AI, Agent Frameworks, Multi-Agent Workflows, AWS, GCP, Microsoft Azure, Observability, Kubernetes, Docker, Terraform, Model Fine-Tuning, Backend Engineering, Full-Stack Engineering
Similar jobs
ML Engineering jobsThe Staff Machine Learning Engineer will architect and deploy scalable generative AI and machine learning systems, including retrieval, inference, evaluation, and agentic workflows. The role requires 7+ years of software development experience, strong Python skills, applied ML expertise, and deep familiarity with modern GenAI platforms and frameworks.
Builds and owns production multi-agent AI infrastructure, backend integrations, and workflow automation for marketing operations. Requires 8+ years of software engineering experience, strong Python and JavaScript/Node.js skills, production LLM experience, and deep Google Cloud expertise.
Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.
Own end-to-end development, evaluation, and production deployment of AI models serving high-volume real-time products. The role requires 5+ years of production Python experience, hands-on fine-tuning and ML operations, cloud infrastructure expertise, and strong technical ownership.
Build and operate production AI agent systems that help Sales and Marketing teams with account planning, deal support, competitive intelligence, and content creation. The role requires 6+ years of experience shipping reliable LLM workflows with retrieval, tool use, permissions, evaluation, and observability.