Build and lead the first ML engineering function at Sprinter Health. Design and implement production ML platforms for training, serving, features, monitoring, retraining and governance; productionize models from prototype to reliable systems in a healthcare startup environment.
220k – 270k/yr
Hybrid8+ YOEML Engineering
About the role
What you will do
Build and lead Sprinter’s ML engineering function as the company’s first dedicated ML engineering hire
Define Sprinter’s ML platform and deployment paradigm across training, serving, features, monitoring, retraining, and governance
Make foundational build-versus-buy, architecture, tooling, and platform decisions that future models and engineers will build on
Design and build production training and inference pipelines that are reliable, observable, and maintainable
Package models for deployment and serve predictions through APIs, batch jobs, or other production workflows
Build clean interfaces between data systems, models, and product systems so ML can be consumed safely and reliably
Maintain feature pipelines and ensure features remain fresh, correct, and consistent between training and serving
Implement monitoring for model performance, drift, data quality, latency, cost, reliability, and production behavior
Prevent training-serving skew, silent degradation, and model regressions before they become production issues
Automate retraining, validation, deployment, rollback, and other production ML workflows where appropriate
Establish reproducibility, versioning, model governance, and operational readiness practices as company defaults
Partner with engineering, data platform, product, operations, and applied science teams to productionize models and improve handoffs
Write design docs, define technical standards, and bring the broader engineering organization along on key ML infrastructure decisions
Set the technical bar for ML engineering by helping interview, mentor, and eventually hire engineers who follow
What you have done
Spent 8+ years building production software, data systems, ML systems, platform infrastructure, or related technical systems
Built and owned ML systems in production across training, serving, features, monitoring, and deployment
Taken models from prototype or research stage into reliable, production-grade systems
Built or meaningfully scaled ML infrastructure, MLOps platforms, model-serving systems, feature pipelines, or related infrastructure
Designed systems that other engineers, data scientists, analysts, or product teams rely on
Made architectural decisions around ML platform design, serving patterns, feature infrastructure, build versus buy, and operational standards
Worked with cloud infrastructure, containers, CI/CD, orchestration, data pipelines, and production deployment workflows
Built monitoring, observability, validation, or alerting for ML systems, data systems, or high-reliability production services
Created reproducible workflows across data, features, models, training runs, deployments, or experiments
Partnered closely with data science, applied science, data platform, product, operations, or backend engineering teams
Operated in ambiguous environments where there was no existing playbook and technical decisions had a long half-life
Balanced speed, simplicity, reliability, privacy, and long-term maintainability in production systems
What gives you an edge
You’ve been an early ML engineer, founding ML engineer, or first ML infrastructure hire at a startup
You’ve built ML infrastructure in a high-growth or operationally complex environment
You have depth in large-scale model serving, feature infrastructure, LLM infrastructure, or real-time inference systems
You have a background in backend engineering, data engineering, MLOps, platform engineering, or infrastructure engineering
You have experience with feature stores, feature pipelines, or production data systems at scale
You’ve helped interview, hire, mentor, or set the technical bar for ML engineers, platform engineers, or data engineers
You’ve worked with healthcare data, PHI, HIPAA-aware systems, or other sensitive data environments
You have experience with security, privacy, governance, or compliance considerations for production ML systems
Viam is seeking a Lead Software Engineer to lead the Data/ML team, owning technical direction, architecture, and delivery. This hands-on role involves managing a team of 5+ engineers, writing code, and driving the reliability and performance of ML training and inference infrastructure.
220k – 250k/yr
HybridML Engineering
Senior Software Engineer, ML/AI Platform
AttentiveUnited States
Senior Software Engineer on the ML Platform team building and operating data, tooling, serving, and inference layers for a PB-scale feature store and ML lifecycle at a high-growth AI marketing company.
220k – 260k/yr
Remote5+ YOEML Engineering
Senior Software Engineer, Machine Learning (Ads)
DiscordSan Francisco, CA
Develops and deploys ML models for ads targeting, ranking, and measurement on Discord's Ads platform. Requires 5+ years ML experience, 3+ years in Ads ML, Python proficiency, and PyTorch/TensorFlow expertise.
220k – 247k/yr
On-site5+ YOEML Engineering
Senior Machine Learning Engineer
Bluesky SocialSeattle, WA
Designs, develops, and maintains ML systems for content recommendations, search, spam detection, and labeling using Bluesky's social graph. Requires 3+ years in ML/data science focused on recommendations/search, Python/PyTorch proficiency, and rapid experimentation skills.
221k – 405k/yr
Remote3+ YOEML Engineering
Senior Machine Learning Engineer, Safety
RedditUnited States
Build, train, and optimize large language models and ML systems to scalably enforce Reddit's community rules and keep users safe. Requires 5+ years MLE experience with NLP, deep learning frameworks, and distributed training.