Skip to content

Principal ML System Engineer

Leads the technical vision, architecture, and roadmap for a company-wide machine learning platform supporting model training, deployment, serving, monitoring, and generative AI. The role requires expert Python and Java skills, large-scale MLOps experience, cloud and Kubernetes expertise, and organization-wide technical leadership.

About the job

Responsibilities

  • Partner with product and engineering leadership to translate business and product objectives into a multi-quarter technical strategy and roadmap for the ML platform.
  • Define reference architectures and standards for scalable data and ML pipelines spanning model training, evaluation, deployment, and serving.
  • Set company-wide MLOps direction and best practices, including model CI/CD, model registries, feature stores, experiment tracking, and build-vs-buy decisions.
  • Establish reliability, observability, and performance practices for production ML systems, including monitoring, alerting, and automated remediation.
  • Establish ML platform security architecture, including authentication, role-based access control, audit logging, and compliance monitoring.
  • Define secure, cost-efficient integration and infrastructure patterns for connecting the platform with existing systems, APIs, and data sources at scale.
  • Provide technical leadership and mentorship across engineering teams and influence the organization-wide ML infrastructure roadmap.

Requirements

  • Expert-level proficiency in Python and Java with strong software engineering fundamentals.
  • Deep experience designing and building ML platforms and MLOps workflows at scale.
  • Experience with cloud platforms, containerization, and orchestration.
  • Demonstrated success setting technical direction and driving organization-wide technical initiatives across multiple teams.

Nice-to-haves

  • Bachelor’s degree or higher in Computer Science, Machine Learning, or a related field.
  • Familiarity with Azure Machine Learning, Databricks processing and serverless environments, and ML frameworks.
  • Experience leading and sustaining critical cross-team systems.
  • Experience implementing security at scale, including role-based access control, multifactor authentication, network security, and compliance monitoring.
  • Experience optimizing large-model training and inference, including LLM serving, for performance and cost.

Compensation

  • Annual salary: $176,000–$195,000.

Skills

Python, Java, Ml Platforms, MLOps, MLflow, Kubeflow, Ray, AWS, Azure, GCP, Docker, Kubernetes, Databricks, Azure Machine Learning, Llm Serving

PointClickCare

PointClickCare

Mississauga, Canada

Principal AI Engineer
CA$192k+/yrHybrid7+ YOEML Engineering

Leads the architecture, development, and governance of agentic AI platforms and reusable patterns across SaaS products. The role requires principal-level experience with Python, Java, Azure and Anthropic AI technologies, RAG pipelines, model serving, evaluation frameworks, and production governance.

Okta

Okta

Toronto, Canada

Staff Machine Learning Engineer, Generative AI
CA$168k+/yrHybrid7+ YOEML Engineering

The Staff Machine Learning Engineer will architect and deploy scalable generative AI and machine learning systems, including retrieval, inference, evaluation, and agentic workflows. The role requires 7+ years of software development experience, strong Python skills, applied ML expertise, and deep familiarity with modern GenAI platforms and frameworks.

Grafana Labs

Grafana Labs

United States
Staff AI Engineer
CA$164k+/yrRemote8+ YOEML Engineering

Builds and owns production multi-agent AI infrastructure, backend integrations, and workflow automation for marketing operations. Requires 8+ years of software engineering experience, strong Python and JavaScript/Node.js skills, production LLM experience, and deep Google Cloud expertise.

Payabli

Payabli

Remote

Staff Machine Learning Engineer
No salary listedRemote8+ YOEML Engineering

Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.

Stream

Stream

Amsterdam, Netherlands
Staff Backend Engineer – AI
No salary listedHybrid7+ YOEML Engineering

Own end-to-end development, evaluation, and production deployment of AI models serving high-volume real-time products. The role requires 5+ years of production Python experience, hands-on fine-tuning and ML operations, cloud infrastructure expertise, and strong technical ownership.