# Machine Learning Engineer

**Company:** [Sprinter Health](https://hotfix.jobs/companies/sprinter-health)
**Location:** San Francisco, CA, Menlo Park, CA
**Role:** ML Engineering
**Salary:** $220k – $270k/yr
**Experience:** 8+ years
**Skills:** MLOps, model serving, feature pipelines, training pipelines, inference pipelines, model monitoring, model governance, llm infrastructure, feature stores, Cloud Infrastructure, CI/CD, Containers, Kubernetes, Data Pipelines, Observability
**Posted:** 2026-07-20

> Build and lead the first ML engineering function at Sprinter Health. Design and implement production ML platforms for training, serving, features, monitoring, retraining and governance; productionize models from prototype to reliable systems in a healthcare startup environment.

## Job Description

## What you will do
- Build and lead Sprinter’s ML engineering function as the company’s first dedicated ML engineering hire
- Define Sprinter’s ML platform and deployment paradigm across training, serving, features, monitoring, retraining, and governance
- Make foundational build-versus-buy, architecture, tooling, and platform decisions that future models and engineers will build on
- Design and build production training and inference pipelines that are reliable, observable, and maintainable
- Package models for deployment and serve predictions through APIs, batch jobs, or other production workflows
- Build clean interfaces between data systems, models, and product systems so ML can be consumed safely and reliably
- Maintain feature pipelines and ensure features remain fresh, correct, and consistent between training and serving
- Implement monitoring for model performance, drift, data quality, latency, cost, reliability, and production behavior
- Prevent training-serving skew, silent degradation, and model regressions before they become production issues
- Automate retraining, validation, deployment, rollback, and other production ML workflows where appropriate
- Establish reproducibility, versioning, model governance, and operational readiness practices as company defaults
- Partner with engineering, data platform, product, operations, and applied science teams to productionize models and improve handoffs
- Write design docs, define technical standards, and bring the broader engineering organization along on key ML infrastructure decisions
- Set the technical bar for ML engineering by helping interview, mentor, and eventually hire engineers who follow

## What you have done
- Spent 8+ years building production software, data systems, ML systems, platform infrastructure, or related technical systems
- Built and owned ML systems in production across training, serving, features, monitoring, and deployment
- Taken models from prototype or research stage into reliable, production-grade systems
- Built or meaningfully scaled ML infrastructure, MLOps platforms, model-serving systems, feature pipelines, or related infrastructure
- Designed systems that other engineers, data scientists, analysts, or product teams rely on
- Made architectural decisions around ML platform design, serving patterns, feature infrastructure, build versus buy, and operational standards
- Worked with cloud infrastructure, containers, CI/CD, orchestration, data pipelines, and production deployment workflows
- Built monitoring, observability, validation, or alerting for ML systems, data systems, or high-reliability production services
- Created reproducible workflows across data, features, models, training runs, deployments, or experiments
- Partnered closely with data science, applied science, data platform, product, operations, or backend engineering teams
- Operated in ambiguous environments where there was no existing playbook and technical decisions had a long half-life
- Balanced speed, simplicity, reliability, privacy, and long-term maintainability in production systems

## What gives you an edge
- You’ve been an early ML engineer, founding ML engineer, or first ML infrastructure hire at a startup
- You’ve built ML infrastructure in a high-growth or operationally complex environment
- You have depth in large-scale model serving, feature infrastructure, LLM infrastructure, or real-time inference systems
- You have a background in backend engineering, data engineering, MLOps, platform engineering, or infrastructure engineering
- You have experience with feature stores, feature pipelines, or production data systems at scale
- You’ve helped interview, hire, mentor, or set the technical bar for ML engineers, platform engineers, or data engineers
- You’ve worked with healthcare data, PHI, HIPAA-aware systems, or other sensitive data environments
- You have experience with security, privacy, governance, or compliance considerations for production ML systems

## Similar roles

- [Senior Software Engineer, Machine Learning](https://hotfix.jobs/jobs/12f53357-b6db-4647-9a5a-2f9e315a3537) - Discord - San Francisco, CA - $220k – $275k/yr
- [Lead Software Engineer, Data Platform](https://hotfix.jobs/jobs/971c7657-3720-45b5-895f-92b209c5be4c) - Viam - New York, NY - $220k – $250k/yr
- [Senior Software Engineer, ML/AI Platform](https://hotfix.jobs/jobs/1b7a2837-b605-4dd4-b03e-aff6f8c85440) - Attentive - Remote - $220k – $260k/yr
- [Senior Machine Learning Engineer](https://hotfix.jobs/jobs/6fc6255d-fb31-4ef8-9926-6151402c0eaf) - Bluesky Social - Remote - $221k – $405k/yr
- [Senior Machine Learning Engineer, Safety](https://hotfix.jobs/jobs/524d6e69-0f3b-4102-bf70-89a599793d6c) - Reddit - Remote - $217k – $303k/yr

**Apply:** https://hotfix.jobs/jobs/4b6250a2-4a7b-49d5-8892-b7f17ff221e9
**Canonical:** https://hotfix.jobs/jobs/4b6250a2-4a7b-49d5-8892-b7f17ff221e9