# Software Engineer, MLOps - Machine Learning

**Company:** [Baton](https://hotfix.jobs/companies/baton)
**Location:** San Francisco, CA
**Role:** ML Engineering
**Salary:** $162k – $216k/yr
**Experience:** 5+ years
**Skills:** Python, MLOps, Distributed Systems, Machine Learning, SQL, Kubernetes, AWS, kubeflow, apache iceberg, feast, SageMaker, batch processing, model monitoring, experiment tracking, Caching
**Posted:** 2026-08-03

> Build and operate production machine-learning infrastructure, automate model lifecycle workflows, and productionize models across distributed systems. The role requires advanced Python, distributed computing, data engineering, SQL, and hands-on MLOps experience, with Kubernetes and cloud infrastructure experience preferred.

## Job Description

## Responsibilities

### Build and Expand MLOps Infrastructure
- Build automated capabilities for model monitoring, retraining, redeployment, champion/challenger testing, A/B testing, and drift detection.
- Improve experiment tracking and model lifecycle management as the number of production models increases.

### Develop and Productionize Machine-Learning Models
- Bring new machine-learning models into production, including developing select models from initial concept through deployment.
- Support models across development, deployment, monitoring, maintenance, and iteration.
- Build scalable batch-prediction capabilities alongside real-time machine-learning workflows.

### Create Self-Serving ML Infrastructure
- Build on existing infrastructure patterns and templates to create reliable and reusable ML workflows.
- Make it easier for engineers to ship and maintain models end to end with less manual intervention.
- Improve development velocity while maintaining production reliability and operational quality.

### Strengthen Distributed ML Systems
- Design and maintain distributed systems that support data-intensive and machine-learning workloads.
- Improve the scalability, performance, and reliability of production ML infrastructure.
- Contribute to batch processing, caching, data movement, and cloud-native infrastructure.

### Connect ML Systems with the Core Platform
- Strengthen the integration between the ML platform and the core transportation management platform.
- Replace manual integration workflows with scalable and maintainable infrastructure.
- Enable machine-learning capabilities to support transportation workflows and operational decision-making.

### Collaborate Across the ML Lifecycle
- Partner with engineers and cross-functional stakeholders to identify opportunities for automation and model productionization.
- Contribute across software engineering, ML development, infrastructure, and production operations based on team needs.

## Requirements

- Advanced proficiency coding production-grade Python at an L4 or L5 level.
- Experience working in an environment where production code directly impacts operations.
- Ability to build and maintain reliable software across modeling, infrastructure, and automation workflows.
- Strong background in distributed computing, scalable ML infrastructure, and high-performance engineering.
- Experience building or maintaining systems that support data-intensive and machine-learning workloads.
- Familiarity with big-data systems, batch processing, caching, and cloud infrastructure.
- Experience implementing, deploying, and productionizing machine-learning algorithms.
- Hands-on experience with data engineering, distributed training, model monitoring, and experiment tracking.
- Experience with model retraining, redeployment, serving, and lifecycle management.
- Strong SQL knowledge and caching experience.

## Nice-to-Haves

- Experience implementing, deploying, monitoring, and maintaining machine-learning models in production.
- Experience with Kubernetes and cloud infrastructure, preferably AWS.
- Familiarity with Kubeflow, Iceberg, Feast, or SageMaker.
- Experience with batch prediction, model serving, distributed training, experiment tracking, caching, or feature stores.
- Experience building scalable, self-serving infrastructure for machine-learning teams.
- Experience integrating ML platforms with broader production or operational systems.
- Experience in a technically rigorous environment such as a large-scale technology company, infrastructure organization, or high-growth engineering team.
- Experience in logistics, transportation, freight, or supply chain.

## Compensation and Benefits

- Annual base salary range: **$162,000–$216,000**.
- Annual company bonus and cash bonus structure.
- Long-term incentive plan.
- 401(k) with matching.
- Hybrid work schedule.
- Medical, dental, and vision coverage.
- Employee stock purchase program with a 15% discount to market value.
- Collaborative, tech-forward office in Hayes Valley, San Francisco.

## Similar roles

- [Software Engineer, ML Infrastructure](https://hotfix.jobs/jobs/4422daeb-5b27-4a07-a336-2e74eac72d8d) - Nuro - Mountain View, CA - $160k – $241k/yr
- [Software Engineer, AI Product Engineer](https://hotfix.jobs/jobs/eed899e1-5bb5-4f1a-8ce5-3e622f1f3044) - Retool - San Francisco, CA - $164k – $306k/yr
- [Applied Scientist, Optimization & Logistics](https://hotfix.jobs/jobs/37773bdb-e364-4a7b-bc8f-3490073e0b9c) - Sprinter Health - San Francisco, CA - $160k – $220k/yr
- [Research Scientist II](https://hotfix.jobs/jobs/4ac4da1a-a497-413d-b8fa-6b3b58cef7e1) - Pindrop - Remote - $160k – $185k/yr
- [AI Engineer - Database Engineering](https://hotfix.jobs/jobs/e2167614-4d69-4c69-bb69-40e1c6ba4571) - Snowflake - Menlo Park, CA - $160k – $230k/yr

**Apply:** https://hotfix.jobs/jobs/d9aa5cd9-c429-4275-9b08-ef42ece4c199
**Canonical:** https://hotfix.jobs/jobs/d9aa5cd9-c429-4275-9b08-ef42ece4c199