Skip to content
Scale AIScale AI

Senior Machine Learning Engineer, Agent Oversight

Senior ML Engineer building observability, evaluation frameworks, and improvement loops for production agentic AI systems. Requires 5+ years production ML/LLM experience, strong grounding in agent design or evaluation, and hands-on work taking systems from prototype to scale.

About the job

Responsibilities

  • Build or contribute to observability into agent behavior in production — the signals and instrumentation needed to actually see what an agent is doing, not just whether it succeeded or failed.
  • Design evaluation methodologies and metrics for agentic applications, and work with the platform to make them run automatically, at scale, across different customer use cases.
  • Build, ship, and own ML systems that detect drift, anomalies, or misalignment in production agent behavior — from first prototype through running reliably at scale.
  • Design and run rigorous experiments to validate model and agent performance improvements before they ship.
  • Work alongside software engineers on the platform where your work intersects with broader infrastructure — take your own work from idea to production.
  • Collaborate closely with product managers, customers, data annotators, Forward Deployed Engineers, and other engineering teams to translate enterprise and government requirements into robust platform capabilities.
  • Depending on focus, contribute to novel methods and approaches that push the state of the art for agent evaluation and improvement, or focus on building ML systems that hold up reliably at scale in production.

Requirements

  • 5+ years of experience as an ML engineer or applied scientist, ideally on a production ML or LLM-powered system — not just consuming a third-party ML API within a feature.
  • Strong grounding in at least two of the following:
    • Building or scaling evaluation, monitoring, or continuous-learning infrastructure for ML/agentic systems.
    • Design experience for agent systems (architecture, orchestration, tool use).
    • Developing new methods, reward models, or model training/fine-tuning approaches.
  • Hands-on experience with LLMs and agent architectures — tool use, planning, multi-agent orchestration.
  • Comfortable partnering with software engineers to productionize research and experimental work, not just deliver a one-off analysis.
  • Rigorous approach to experimentation: clear hypotheses, real statistical grounding, and results that hold up under scrutiny.
  • Track record of collaborating across functions (Product, Forward Deployed Engineering, etc.) to navigate ambiguous requirements and bring them to production.
  • Gives direct, substantive feedback on designs and code, and takes it the same way — and mentors others as they grow.

Nice-to-Haves

  • Experience building or contributing to RLHF, SFT, or other fine-tuning/RL workflows, reward modeling, or verifiable-reward systems.
  • Experience with model or systems optimization (e.g., latency, cost, or inference efficiency).
  • Published research, open-source contributions, or patents in agentic systems, LLMs, or applied ML.
  • Experience working in regulated or enterprise contexts.
  • Track record of taking a novel method from prototype to something running reliably in production, navigating ambiguity along the way.
  • Experience reviewing others’ technical designs or mentoring engineers at a senior/staff level.

Skills

LLMs, Agent Architectures, Evaluation Frameworks, Observability Tools, RLHF, Sft, Reward Modeling, Fine-Tuning, Ml Experimentation, Python

Reddit

Reddit

Ontario, Canada

Senior Machine Learning Engineer, Ads
$217k+/yrRemote5+ YOEML Engineering

Design, build, and deploy production ML systems for recommendations, search, ranking, and advertising at internet scale. Own the full ML lifecycle from modeling to monitoring with strong cross-functional collaboration.

Fetch

Fetch

United States

Senior Machine Learning Engineer II
$211k+/yrRemote6+ YOEML Engineering

Build and operate low-latency machine learning systems for ad ranking, relevance, and optimization, including feature pipelines, experimentation, evaluation, and production inference. The role requires 6+ years of software engineering experience, strong Python skills, AWS experience, and practical LLM application experience.

Dialpad

Dialpad

United States

Senior AI Engineer
$225k+/yrRemote5+ YOEML Engineering

Leads development of speech models, decoders, and low-latency inference systems for next-generation voice agents. Requires 5+ years in speech ML or related audio AI, strong Python and PyTorch experience, and the ability to guide technical direction and mentor engineers.

Checkr

Checkr

San Francisco, CA

Senior Machine Learning Engineer
$207k+/yrOn-site6+ YOEML Engineering

Build and operate production ML and AI services using Python, LLM APIs, and robust software engineering practices. The role requires 6+ years of professional software experience, including production ML systems, and partners closely with product and engineering teams.

Ambience Healthcare

Ambience Healthcare

San Francisco, CA

Senior Machine Learning Engineer
$225k+/yrHybrid5+ YOEML Engineering

Build and improve production AI systems for clinical products, owning evaluations, model behavior, agentic workflows, data flywheels, deployment, and observability. The role requires 5+ years of production ML or applied AI experience, strong Python and modern ML framework skills, and hands-on debugging expertise.