Senior Machine Learning Operations Engineer
Build and operate the platform that deploys, serves, observes, and retrains production machine-learning models for real-time fraud and financial-crime risk decisions. Requires 5+ years of ML engineering, backend, or MLOps experience, strong Python skills, and production model-serving expertise.
About the job
Responsibilities
- Build and operate real-time inference services that score models for risk decision engines, prioritizing low latency and high availability.
- Own model deployment infrastructure, including registries, versioning, model CI/CD, performance, bias and consistency checks, shadow mode, and staged rollouts.
- Build production model observability for availability, latency, errors, and drift detection that can trigger retraining.
- Partner with Risk Data Science to move models from development into production and operate them through the Machine Learning Platform.
- Implement experimentation capabilities such as champion/challenger and canary routing, as well as explainability outputs such as SHAP attributions.
- Help shape and build a new machine learning platform team with strong product ownership.
Requirements
- 5+ years of experience in machine learning engineering, backend software engineering, MLOps, or a closely related field.
- Experience deploying, serving, and operating production ML models in low-latency, highly available environments.
- Strong backend engineering fundamentals in Python, with experience using API frameworks such as FastAPI or Flask.
- Experience with model deployment and lifecycle tooling, including model registries, model CI/CD, versioning, and staged rollout patterns.
- Experience building observability and alerting for production services, including latency and error monitoring; model-specific signals such as drift are preferred.
- Familiarity with SQL, key-value or low-latency data stores, and streaming pipelines.
Nice to Have
- Familiarity with Snowflake, dbt, Dagster, Airflow, or similar modern data-stack tools.
- Experience in regulated, audit-sensitive, or compliance-adjacent environments.
- Exposure to functional languages or willingness to work with Haskell, React, and TypeScript.
Compensation
- US employees: $166,600–$208,300 USD base salary.
- Canadian employees: $157,400–$196,800 CAD base salary.
- Total rewards include base salary, equity, and benefits.
Skills
Machine Learning, MLOps, Python, FastAPI, Flask, Model Registries, CI/CD, Shap, SQL, Redis, DynamoDB, Kafka, Snowflake, Airflow, TypeScript
Similar jobs
ML Engineering jobsLeads hands-on development and deployment of production AI agents and workflows, sets technical standards, and mentors engineers. Requires 4+ years of software engineering experience, production LLM application experience, and strong Python or TypeScript skills.
Owns the full lifecycle of data and ML solutions, from ingestion and feature-ready datasets through production deployment and business-impact measurement. The role combines data engineering, applied machine learning, MLOps, and generative AI to build risk detection capabilities.
Develop and deploy real-time perception and sensor-fusion software for autonomous battery-electric rail vehicles. The role requires strong robotics, geometry-based computer vision, C/C++ and Rust experience, plus hands-on work with multimodal sensors and production systems.
Build and deploy production machine-learning models and data systems that classify and enrich Internet telemetry for internal platforms and customer-facing products. The role requires 5+ years of applied ML, data science, or software engineering experience, plus strong Python or Go skills.
Designs and ships production multi-agent compliance systems, including LLM pipelines, model training, evaluation, monitoring, and explainability. Requires 5+ years of applied AI/ML engineering experience, strong Python, and experience deploying production ML systems.