Staff Machine Learning Operations Engineer
Leads the reliability, architecture, deployment automation, and monitoring of production machine learning systems. Requires 7+ years of software engineering experience, deep MLOps platform expertise, and strong Kubernetes, cloud, infrastructure-as-code, and observability fundamentals.
About the job
Responsibilities
- Own the reliability, performance, functionality, and cost-efficiency of production machine learning systems, including SLOs, observability, and on-call responsibilities.
- Architect the ML platform, including a feature store, model registry, ML CI/CD, data infrastructure, and standardized service patterns.
- Build automated, pull-request-driven ML CI/CD workflows with data quality checks and statistical model validation before deployment.
- Improve cost and latency through architecture, hardware, and model optimization.
- Establish workflows, standards, KPIs, and onboarding practices for a future MLOps team.
- Design automated data drift and concept drift monitoring, creating alerts for model degradation and preparing for continuous training architectures.
- Collaborate with ML, data, platform, data science, and product engineering teams; set technical direction for MLOps.
Requirements
- 7+ years of software engineering experience, including substantial experience operating ML or data-intensive systems in production at scale.
- Deep experience with model serving, feature stores, model registries, and CI/CD for machine learning.
- Strong infrastructure and platform engineering fundamentals, including Kubernetes, containers, cloud infrastructure, Terraform/IaC, observability, and incident response.
- Experience designing ML platforms or significant platform components, with judgment about when to build versus buy.
Nice-to-haves
- Experience with healthcare, regulated data, or other high-stakes production ML systems.
Technologies
- Python
- Kubernetes
- AWS
- Amazon SageMaker
- Terraform
- Amazon S3
- Snowflake
- Apache Airflow
- Datadog
Compensation
- Target salary range: $298,000–$351,000 annually.
- Eligible for equity incentives and benefits including flexible PTO, medical, dental, vision, 401(k), and Teladoc Health.
Skills
Python, Kubernetes, AWS, Amazon Sagemaker, Terraform, Amazon S3, Snowflake, Apache Airflow, Datadog, Feature Stores, Model Serving, Model Registries, CI/CD
Similar jobs
ML Engineering jobsLeads end-to-end development of production algorithmic systems for healthcare, spanning machine learning, optimization, and LLM applications. The player-coach role requires 6+ years of industry experience, strong problem-solving and metrics judgment, and technical leadership of a small team.
Leads technical strategy for Reddit’s Ads ML Platform, improving feature development, training-data generation, experimentation, and the path to production ML serving. The role requires 8+ years in infrastructure or distributed systems, production ML platform experience, and strong cross-team technical leadership.
Build and operate production machine learning systems for ranking, retrieval, recommendations, personalization, and customer intelligence. The role requires 12+ years of production software and ML experience, strong expertise in intelligent systems, and sound judgment around trustworthy customer-impacting signals.
Build and operate scalable ML inference infrastructure for Claude’s safety systems, translating safety research into reliable production deployments. The role requires deep production ML infrastructure experience, distributed systems expertise, and proficiency with Python and modern ML frameworks.
Leads the technical direction and development of large-scale, GenAI-powered recommendation and feed-ranking systems. Requires 10+ years of industry experience in relevance-driven products, deep expertise in machine learning and recommendations, and strong organizational influence and mentoring skills.