Senior Machine Learning Engineer
Build and improve production AI systems for clinical products, owning evaluations, model behavior, agentic workflows, data flywheels, deployment, and observability. The role requires 5+ years of production ML or applied AI experience, strong Python and modern ML framework skills, and hands-on debugging expertise.
About the job
Responsibilities
- Design and own evaluation pipelines for LLM and agentic systems using automated graders, regression testing, production feedback, and human evaluation.
- Diagnose production failure modes and improve model behavior through prompting, retrieval, context, routing, data, fine-tuning, and other system interventions.
- Build production agentic AI systems with tool use, retrieval, context and state management, routing, orchestration, tracing, and failure recovery.
- Create datasets, evaluations, and feedback loops using production failures, user feedback, and active learning.
- Translate research in LLMs, agents, NLP, speech, and multimodal AI into practical experiments.
- Own AI systems end-to-end across models, data, evaluation, orchestration, serving, and observability.
- Collaborate with clinicians, product managers, and engineers; communicate complex AI concepts clearly.
Requirements
- 5+ years of experience in production machine learning, research engineering, or applied AI.
- Experience building a consequential production AI system or materially improving model behavior in production.
- Strong understanding of modern LLMs, transformers, and production AI systems.
- Experience designing evaluations for LLMs, agents, or other complex AI systems.
- Ability to turn ambiguous quality problems into measurable dimensions, datasets, and experiments.
- Familiarity with grader bias, leakage, misleading aggregate metrics, regression detection, and offline-online mismatch.
- Experience building production systems involving multiple models, tools, retrieval, context, state, routing, or orchestration.
- Proficiency in Python and modern machine learning frameworks; PyTorch preferred.
- Comfort with deployment, observability, CI/CD, and containerized systems.
- Hands-on experience writing code, inspecting traces, analyzing failures, and debugging production systems.
- Experience building high-quality datasets and feedback loops.
- Strong interdisciplinary collaboration, communication, ownership, and problem-solving skills.
Nice-to-Haves
- Experience with real-time voice, conversational AI, or multimodal systems.
- Experience with fine-tuning, post-training, or model adaptation.
- Experience in healthcare, clinical AI, or other regulated, high-stakes industries.
- Experience interviewing or mentoring machine learning engineers.
- Open-source contributions to machine learning, agent, or evaluation tooling.
Compensation and Benefits
- $225,000–$300,000, plus significant equity.
- Comprehensive medical, dental, and vision coverage for employees and dependents.
- 401(k) with company match of up to 3% of base salary.
- Remote-friendly culture with a San Francisco headquarters and equipment provisioning.
- Parental leave.
- Annual company-wide and team off-sites, team lunches, and all-hands gatherings with travel, lodging, and meals covered.
- Flexible time off with no annual cap, company holidays, and an annual holiday shutdown from December 24–January 1.
Skills
Python, PyTorch, LLMs, Transformers, Machine Learning, NLP, Retrieval, Fine-Tuning, Active Learning, CI/CD, Docker, Observability, Agentic AI, Multimodal Ai, Conversational AI
Similar jobs
ML Engineering jobsLeads development of speech models, decoders, and low-latency inference systems for next-generation voice agents. Requires 5+ years in speech ML or related audio AI, strong Python and PyTorch experience, and the ability to guide technical direction and mentor engineers.
Senior AI Engineer responsible for production LLM agents that enrich business identity data through web discovery, verification, classification, and risk scoring. The role requires strong asynchronous Python, agent and evaluation expertise, browser automation, and experience operating AI systems in production.
Leads the development and production deployment of large-scale ASR and TTS systems for conversational intelligence products. The role requires 5+ years of industry experience, deep speech-model expertise, and strong software engineering and ML operations capabilities.
Design, build, and deploy production ML systems for recommendations, search, ranking, and advertising at internet scale. Own the full ML lifecycle from modeling to monitoring with strong cross-functional collaboration.
Build and operate low-latency machine learning systems for ad ranking, relevance, and optimization, including feature pipelines, experimentation, evaluation, and production inference. The role requires 6+ years of software engineering experience, strong Python skills, AWS experience, and practical LLM application experience.