Build and productionize agentic LLM workflows for clinical documentation at Abridge, owning architecture, rigorous evaluation frameworks, and deployment of reliable GenAI systems in healthcare. Requires 3+ years building production systems with 1-2 years on LLMs or agentic products, strong Python skills, and hybrid work in SF or NYC.
255k – 300k/yr
Hybrid3+ YOEML Engineering
About the role
What You’ll Do
Design and build agentic systems that turn LLMs into composable, dependable tools—leveraging retrieval, tool use, agentic reasoning, and structured outputs.
Collaborate with ML and infra engineers to scale and optimize agentic workflows, managing latency, context windows, and model choice.
Write high-quality, modular code that’s graceful under failure, flexible to change, and easy to iterate on.
Own major architectural decisions—how we architect workflows, define data flow, cache intermediate state, and structure generative outputs.
Drive rigorous evaluation: build benchmark datasets, develop automated and human-in-the-loop frameworks, design experiments to surface failure modes and edge cases, run A/B tests to inform deployment, and distill insights from clinician feedback to evaluate and guide model improvement.
Leverage frontier capabilities: rapidly prototype with new models and model capabilities, open-source tools, and novel prompting techniques.
What You’ll Bring
3+ years of experience building production-grade systems, with 1–2+ years focused on LLM-powered or agentic products.
Deep fluency with LLM APIs, prompting strategies, and orchestration patterns (e.g., LangChain, LlamaIndex, custom pipelines).
Experience with retrieval systems (e.g., semantic and lexical retrieval, vector DBs, efficient kNN), function calling, tool-use, or agentic workflows.
Working knowledge of model evaluation, experience building diverse datasets, conducting both automated and human-in-the-loop evaluations, running A/B tests, and working with subject matter experts to guide model improvement.
Strong Python fundamentals—including ability to write clean code, design comprehensive test-cases, and familiarity with core language features and standard libraries; experience with async programming, performance profiling, packaging, and deployment tooling is strongly preferred.
Good taste and intuition: You know when to move fast, ship, and iterate and also when to take a beat to tackle tech debt.
We value people who are eager to learn new things and recognize that great team members might not perfectly match a job description. If you’re interested in the role but aren’t sure whether or not you’re a good fit, we’d still like to hear from you.
Must be willing to work from our SF or NYC office at least 3x per week.
Compensation and Benefits
Competitive compensation and equity grants for full time employees.
Generous Time Off: 14 paid holidays, flexible PTO for salaried employees, and accrued time off for hourly employees.
Comprehensive Health Plans: Medical, Dental, and Vision coverage for all full-time employees and their families.
Generous HSA Contribution: If you choose a High Deductible Health Plan, Abridge makes monthly contributions to your HSA.
Paid Parental Leave: Generous paid parental leave for all full-time employees.
Family Forming Benefits: Resources and financial support to help you build your family.
401(k) Matching: Contribution matching to help invest in your future.
Personal Device Allowance: Tax free funds for personal device usage.
Pre-tax Benefits: Access to Flexible Spending Accounts (FSA) and Commuter Benefits.
Lifestyle Wallet: Monthly contributions for fitness, professional development, coworking, and more.
Mental Health Support: Dedicated access to therapy and coaching to help you reach your goals.
Sabbatical Leave: Paid Sabbatical Leave after 5 years of employment.
Builds core infrastructure for AI software agents in Devin and Windsurf, focusing on long-horizon task execution, tool use, planning, and reliability at scale. Requires strong Python, systems engineering depth, and AI curiosity; onsite in San Francisco.
260k – 300k/yrOn-siteML Engineering
Applied AI Inference Engineer
CrusoeSan Francisco, CA +1
Build and optimize the end-to-end LLM inference stack for production deployments. Profile and tune serving frameworks (vLLM/SGLang) and CUDA kernels for latency, throughput and cost; partner directly with customer teams to take workloads from POC to monitored production.
250k – 300k/yrOn-site5+ YOEML Engineering
Research Scientist / Engineer — Multimodal Agent
Luma AIPalo Alto, CA
Builds and trains large-scale multimodal agentic models involving reasoning, planning, coding, and tool calling. Requires strong ML foundations, PyTorch expertise, and experience with distributed training on massive datasets.
250k – 450k/yrHybridML Engineering
Research Engineer, Evals
VarianceSan Francisco, CA
Build benchmarks, datasets, and evaluation systems to measure and improve AI model quality for fraud, identity, and risk judgment tasks. Collaborate across research, engineering, and product to drive rigorous experimentation and iteration in high-stakes environments.
250k – 400k/yrOn-siteML Engineering
Research Engineer, Judgment Systems
VarianceSan Francisco, CA
Research Engineer designs evaluations, studies model failures, and builds research loops to improve AI agents for high-stakes fraud detection and judgment tasks. Requires ML training experience, experimental rigor, and strong engineering skills in adversarial environments.