Skip to content

Principal AI Engineer

Leads the architecture, development, and governance of agentic AI platforms and reusable patterns across SaaS products. The role requires principal-level experience with Python, Java, Azure and Anthropic AI technologies, RAG pipelines, model serving, evaluation frameworks, and production governance.

About the job

Responsibilities

  • Set the technical vision and reference architecture for agentic AI across applications, including standards for reasoning loops, tool use/function calling, memory, safety, and interoperability.
  • Build and govern reusable platform components such as model brokers, multi-model routing, orchestration services, evaluation harnesses, prompt/versioning stores, and telemetry.
  • Develop and refine advanced prompting strategies for large language models.
  • Drive cross-functional roadmaps and integration standards, including API and versioning contracts, LLM and agent cost/performance optimization, and engineering mentorship.
  • Establish evaluation frameworks for LLMs and agentic systems, covering offline and online testing, safety assessments, quality metrics, and business-outcome dashboards.
  • Serve as an internal AI subject-matter expert for enterprise applications, leadership, vendors, business partners, and technical teams.
  • Identify, prioritize, and execute AI use cases with measurable business outcomes.
  • Coach and mentor engineers through peer reviews, knowledge-sharing sessions, engineering best practices, pull-request approvals, and technical-quality improvements.
  • Transition prototypes into secure, observable, production-ready solutions while aligning with responsible AI, security, and regulatory requirements.

Requirements

  • Proven leadership shipping agentic AI at scale across multiple products and establishing organization-wide technical standards.
  • Hands-on expertise with Microsoft AI technologies, including Azure AI Foundry, Azure OpenAI, and Azure AI Document Intelligence.
  • Experience with Anthropic Claude Cowork and Claude Code.
  • Advanced Python and Java expertise in production systems, including testing, packaging, and performance profiling.
  • Deep experience extending agent frameworks such as Anthropic, LangChain, and Foundry.
  • RAG and agent data-pipeline expertise, including document preprocessing, OCR/NER, chunking strategies, embeddings, vector stores, data quality, and versioning.
  • Experience evaluating AI systems for quality, cost, and latency.
  • Extensive knowledge of enabling and maturing AI in a SaaS environment.
  • Experience with model serving and cost/performance optimization, including model-broker patterns, multi-model routing, and serving stacks.
  • Knowledge of relational databases such as Microsoft SQL Server and PostgreSQL.
  • Experience with system and performance monitoring tools such as AppDynamics, Kibana, and OpenTelemetry.
  • BSc/BA in Computer Science or a related field.

Nice-to-haves

  • Docker, Kubernetes, and Istio experience.
  • AWS or Azure cloud-services experience, or equivalent.
  • Service-oriented architecture experience.
  • On-call experience with production-grade systems.

Skills

Python, Java, Azure Ai Foundry, Azure Openai, Azure Ai Document Intelligence, Anthropic Claude, LangChain, RAG, Ocr, Ner, Embeddings, Vector Stores, Kubernetes, OpenTelemetry

PointClickCare

PointClickCare

Mississauga, Canada

Principal ML System Engineer
CA$176k+/yrRemote7+ YOEML Engineering

Leads the technical vision, architecture, and roadmap for a company-wide machine learning platform supporting model training, deployment, serving, monitoring, and generative AI. The role requires expert Python and Java skills, large-scale MLOps experience, cloud and Kubernetes expertise, and organization-wide technical leadership.

Okta

Okta

Toronto, Canada

Staff Machine Learning Engineer, Generative AI
CA$168k+/yrHybrid7+ YOEML Engineering

The Staff Machine Learning Engineer will architect and deploy scalable generative AI and machine learning systems, including retrieval, inference, evaluation, and agentic workflows. The role requires 7+ years of software development experience, strong Python skills, applied ML expertise, and deep familiarity with modern GenAI platforms and frameworks.

Grafana Labs

Grafana Labs

United States
Staff AI Engineer
CA$164k+/yrRemote8+ YOEML Engineering

Builds and owns production multi-agent AI infrastructure, backend integrations, and workflow automation for marketing operations. Requires 8+ years of software engineering experience, strong Python and JavaScript/Node.js skills, production LLM experience, and deep Google Cloud expertise.

Payabli

Payabli

Remote

Staff Machine Learning Engineer
No salary listedRemote8+ YOEML Engineering

Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.

Stream

Stream

Amsterdam, Netherlands
Staff Backend Engineer – AI
No salary listedHybrid7+ YOEML Engineering

Own end-to-end development, evaluation, and production deployment of AI models serving high-volume real-time products. The role requires 5+ years of production Python experience, hands-on fine-tuning and ML operations, cloud infrastructure expertise, and strong technical ownership.