Skip to content
BatonBatonSan Francisco, CA

Software Engineer, MLOps - Machine Learning

Build and operate production machine-learning infrastructure, automate model lifecycle workflows, and productionize models across distributed systems. The role requires advanced Python, distributed computing, data engineering, SQL, and hands-on MLOps experience, with Kubernetes and cloud infrastructure experience preferred.

162k – 216k/yr
Hybrid5+ YOEML Engineering

About the role

Responsibilities

Build and Expand MLOps Infrastructure

  • Build automated capabilities for model monitoring, retraining, redeployment, champion/challenger testing, A/B testing, and drift detection.
  • Improve experiment tracking and model lifecycle management as the number of production models increases.

Develop and Productionize Machine-Learning Models

  • Bring new machine-learning models into production, including developing select models from initial concept through deployment.
  • Support models across development, deployment, monitoring, maintenance, and iteration.
  • Build scalable batch-prediction capabilities alongside real-time machine-learning workflows.

Create Self-Serving ML Infrastructure

  • Build on existing infrastructure patterns and templates to create reliable and reusable ML workflows.
  • Make it easier for engineers to ship and maintain models end to end with less manual intervention.
  • Improve development velocity while maintaining production reliability and operational quality.

Strengthen Distributed ML Systems

  • Design and maintain distributed systems that support data-intensive and machine-learning workloads.
  • Improve the scalability, performance, and reliability of production ML infrastructure.
  • Contribute to batch processing, caching, data movement, and cloud-native infrastructure.

Connect ML Systems with the Core Platform

  • Strengthen the integration between the ML platform and the core transportation management platform.
  • Replace manual integration workflows with scalable and maintainable infrastructure.
  • Enable machine-learning capabilities to support transportation workflows and operational decision-making.

Collaborate Across the ML Lifecycle

  • Partner with engineers and cross-functional stakeholders to identify opportunities for automation and model productionization.
  • Contribute across software engineering, ML development, infrastructure, and production operations based on team needs.

Requirements

  • Advanced proficiency coding production-grade Python at an L4 or L5 level.
  • Experience working in an environment where production code directly impacts operations.
  • Ability to build and maintain reliable software across modeling, infrastructure, and automation workflows.
  • Strong background in distributed computing, scalable ML infrastructure, and high-performance engineering.
  • Experience building or maintaining systems that support data-intensive and machine-learning workloads.
  • Familiarity with big-data systems, batch processing, caching, and cloud infrastructure.
  • Experience implementing, deploying, and productionizing machine-learning algorithms.
  • Hands-on experience with data engineering, distributed training, model monitoring, and experiment tracking.
  • Experience with model retraining, redeployment, serving, and lifecycle management.
  • Strong SQL knowledge and caching experience.

Nice-to-Haves

  • Experience implementing, deploying, monitoring, and maintaining machine-learning models in production.
  • Experience with Kubernetes and cloud infrastructure, preferably AWS.
  • Familiarity with Kubeflow, Iceberg, Feast, or SageMaker.
  • Experience with batch prediction, model serving, distributed training, experiment tracking, caching, or feature stores.
  • Experience building scalable, self-serving infrastructure for machine-learning teams.
  • Experience integrating ML platforms with broader production or operational systems.
  • Experience in a technically rigorous environment such as a large-scale technology company, infrastructure organization, or high-growth engineering team.
  • Experience in logistics, transportation, freight, or supply chain.

Compensation and Benefits

  • Annual base salary range: $162,000–$216,000.
  • Annual company bonus and cash bonus structure.
  • Long-term incentive plan.
  • 401(k) with matching.
  • Hybrid work schedule.
  • Medical, dental, and vision coverage.
  • Employee stock purchase program with a 15% discount to market value.
  • Collaborative, tech-forward office in Hayes Valley, San Francisco.

Skills

PythonMLOpsDistributed SystemsMachine LearningSQLKubernetesAWSkubeflowapache icebergfeastSageMakerbatch processingmodel monitoringexperiment trackingCaching

Similar roles

ML Engineering jobs
Nuro

Software Engineer, ML Infrastructure

NuroMountain View, CA

Build and scale ML infrastructure platform for autonomous vehicle development, focusing on automated resource provisioning, high-performance workload scheduling, and petabyte-scale data processing pipelines.

160k – 241k/yrOn-site3+ YOEML Engineering
Retool

Software Engineer, AI Product Engineer

RetoolSan Francisco, CA

Builds AI-powered features for developer tools, including generative UIs from natural language, LLM performance improvements, and RAG patterns. Requires 3+ years experience with LLMs, modern AI techniques, and full-stack engineering in TypeScript/Node.js/React.

164k – 306k/yrHybrid3+ YOEML Engineering
Sprinter Health

Applied Scientist, Optimization & Logistics

Sprinter HealthSan Francisco, CA

Applied Scientist building optimization, forecasting, and simulation models to solve complex logistics and clinician-patient matching problems for in-home healthcare delivery. Requires strong operations research foundations, Python/SQL/ML expertise, and experience shipping production decision systems.

160k – 220k/yrHybridML Engineering
Pindrop

Research Scientist II

PindropUnited States

Research Scientist II building and improving fraud risk models and scam detection systems using audio, behavioral, and metadata signals. Requires an advanced degree and 3+ years of applied ML experience with Python and modern ML frameworks.

160k – 185k/yrRemote3+ YOEML Engineering
Snowflake

AI Engineer - Database Engineering

SnowflakeMenlo Park, CA

As an AI Engineer, you will own the full AI engineering lifecycle, from design to optimization, for Snowflake Database Engineering products. You will build agentic workflows, coding harnesses, and evaluation pipelines, working with a high-powered engineering team.

160k – 230k/yrOn-site5+ YOEML Engineering