Build, evaluate, and productionize ML/AI models (including LLMs and NLP) that solve ambiguous healthcare, product, and operational problems at Sprinter Health. Requires strong experimentation, error analysis, stakeholder collaboration with clinicians, and focus on real-world impact, bias, and evaluation.
180k – 260k/yr
HybridML Engineering
About the role
What you will do
Turn ambiguous healthcare, product, and operational problems into well-posed ML, AI, ranking, optimization, NLP, or LLM-based tasks
Build strong baselines and improve on them efficiently using the right modeling approach for the problem
Develop models across traditional ML, deep learning, NLP, and LLM-based approaches where appropriate
Design offline and online evaluations that are honest, measurable, and predictive of real-world impact
Choose metrics suited to imbalanced, delayed, noisy, and partially observed healthcare outcomes
Run careful error analysis and use it to improve model quality, product fit, and operational usefulness
Identify label leakage, selection bias, confounding, and other data artifacts before they reach production
Explore messy real-world data, assess label quality, and determine whether a problem is ready for modeling
Partner with ML engineering to productionize models reliably and define what production-readiness requires
Work with clinical stakeholders and subject-matter experts to validate assumptions, review model errors, and understand edge cases
Explain model tradeoffs, uncertainty, limitations, and expected impact clearly to product, operations, clinical, and leadership teams
Write experiment docs, summarize findings, and help teams make informed decisions about when and how to deploy AI systems
Pressure-test whether results are real, robust, and useful before recommending production use
What you have done
Built, evaluated, and iterated on machine learning or AI models for real-world use cases
Turned ambiguous business, product, clinical, or operational problems into measurable modeling tasks
Designed rigorous offline evaluations, experiments, or analyses that informed production or product decisions
Worked with messy real-world datasets where labels, outcomes, and causal relationships are imperfect
Used statistical reasoning, experimental design, and error analysis to understand model performance
Built models using Python and standard ML or AI tooling such as PyTorch, scikit-learn, NumPy, pandas, Polars, Hugging Face, Matplotlib, or similar
Compared modeling approaches and made pragmatic decisions about when to use traditional ML, LLMs, heuristics, or simpler baselines
Communicated model performance, limitations, tradeoffs, and uncertainty to technical and non-technical stakeholders
Partnered with engineering, product, data, operations, clinical, or domain experts to move models closer to production impact
Operated with enough engineering depth to run experiments end to end and self-serve deployments or production handoffs when needed
Used AI coding assistants such as Claude Code, Cursor, or similar tools as part of your development workflow
What gives you an edge
You have an MS or PhD in computer science, statistics, machine learning, applied math, operations research, biomedical informatics, epidemiology, or a related quantitative field
You have exceptional applied experience that substitutes for formal graduate training
You have depth in LLMs, ranking, NLP, uncertainty quantification, causal inference, optimization, or healthcare AI
You’ve shipped models that reached production and had measurable real-world impact
You’ve worked with healthcare data such as claims, EHR, clinical notes, scheduling, utilization, quality, risk, or patient engagement data
You have experience working with PHI, HIPAA-aware systems, or other sensitive regulated data
You know when traditional ML approaches are likely to outperform LLMs, and when LLMs are the right tool
You have experience collaborating with clinicians, clinical operations teams, or other high-stakes domain experts
You’ve worked in a startup or fast-moving applied environment where ambiguity, speed, and rigor all mattered
What makes you successful
You understand how ML models work under the hood and can explain them clearly to non-technical stakeholders
You focus relentlessly on impact and know that the simplest model is often the best one
You treat evaluation as one of the most important parts of model development
You notice when a metric is misleading, incomplete, or disconnected from real-world outcomes
You catch leakage, bias, and confounding that others miss
You move fluidly between modeling, error analysis, stakeholder partnership, and production handoff
You can hand a model to engineering and explain its limits to a clinician with equal clarity
You are comfortable with ambiguity and can adapt modeling approaches to problems that do not come with a playbook
You balance scientific rigor with the practical need to ship useful systems
Perception Engineer owning outcomes for autonomous mining vehicles. Responsible for sensor selection, model adaptation to new sites/domains, diagnosing failures, data strategies, and translating customer needs into technical KPIs and solutions. Requires strong systems understanding of perception/full stack and real-world deployment experience.
180k – 255k/yr
On-site5+ YOEML Engineering
Software Development Engineer in Test, Machine Learning
ZooxFoster City, CA
Build and productionize agentic LLM-powered triage systems and ML/DL pipelines to automate failure analysis for autonomous robots. Requires Master's/PhD in STEM + 4+ years production ML/NLP experience with PyTorch, RAG, Databricks, and AWS.
180k – 225k/yr
Hybrid4+ YOEML Engineering
Software Engineer, AI Platform
NotionSan Francisco, CA +1
Build and scale the shared AI platform foundations at Notion, enabling fast and safe shipping of AI products. Requires experience with LLM/ML platforms, strong ownership, and comfort across backend, infrastructure, and product code.
180k – 201k/yr
Hybrid5+ YOEML Engineering
Research Engineer, Generalist
ExaSan Francisco, CA
Generalist Research Engineer working across Exa's search and retrieval stack including crawling, parsing, ML performance, and retrieval algorithms to improve search quality and performance for customers.
180k – 350k/yr
On-siteML Engineering
Software Engineer - BIS
BasetenSan Francisco, CA
As a Software Engineer on the Inference Stack team, you will build the distributed runtime that powers large-scale LLM inference. This role involves working across the stack, from developer experience to low-level infrastructure, and owning systems in production.