# Staff AI Evaluation Lead

**Company:** [Fetch](https://hotfix.jobs/companies/fetch)
**Location:** Remote
**Role:** ML Engineering
**Salary:** $129k – $152k/yr
**Experience:** 8+ years
**Skills:** SQL, JSON, APIs, Python, LLMs, Agentic Systems, Machine Learning, Ai Evaluation, Data Pipelines, Automation Platforms, System Design, Production Systems
**Posted:** 2026-08-07

> Leads organization-wide AI evaluation and automation programs, establishing quality standards, metrics, production gates, and scalable evaluation infrastructure for LLM and agentic systems. The role requires 8+ years of relevant experience, strong technical systems expertise, and cross-functional leadership.

## Job Description

## Responsibilities
- Lead complex, high-impact automation and evaluation initiatives across workflows and teams.
- Architect end-to-end workflows integrating datasets, evaluations, automations, and human-in-the-loop processes.
- Build reusable harnesses, components, and evaluation pipelines that run against production without hands-on operation.
- Define dataset standards, evaluation methodology, failure taxonomies, and quality measurement across AI Operations.
- Create, maintain, and validate metric definitions against real production behavior; revise them as models, tooling, and architectures change.
- Define production quality bars, determine where human review remains in the loop, and remove systems from production when they no longer meet standards.
- Allocate evaluation depth according to system risk and make accepted tradeoffs visible to stakeholders.
- Define when model or platform changes require organizational re-baselining and manage that process.
- Evaluate new models and AI capabilities and determine adoption decisions.
- Redesign workflows, tooling, and processes to improve performance and durability across AI Operations.
- Partner with Engineering, Product, and AI teams to align priorities and deliver solutions.
- Define success metrics, communicate performance and recommendations to leadership, and drive measurable business impact.
- Mentor technical contributors and elevate the team’s automation and evaluation capabilities.
- Stay current on emerging tools and approaches, piloting and translating them into actionable improvements.

## Requirements
- **8+ years** of professional experience in AI, machine learning, operational automation, or a related field.
- Proven ability to lead complex automation or evaluation initiatives across systems or teams.
- Experience designing, building, and scaling evaluation frameworks, datasets, and quality systems for LLMs or AI products.
- Fluency in SQL, JSON, APIs, and scripting with AI assistance, with ownership of the technical direction of an evaluation stack.
- Deep working knowledge of LLM and agentic system behavior, automation platforms, and system design, demonstrated through systems built.
- Experience with data pipelines, APIs, and production systems.
- Experience influencing cross-functional stakeholders and aligning priorities.
- Demonstrated ability to define metrics and drive measurable business impact.

## Preferred Qualifications
- Experience mentoring or leading technical contributors.
- Experience evaluating agentic systems in production at scale.
- Experience setting technical standards adopted across an organization.

## Compensation and Benefits
- Base salary range: **$129,000 - $152,000**.
- Equity for full-time employees.
- Dollar-for-dollar 401(k) match up to 4%.
- Medical, dental, and vision plans, including pet coverage.
- Up to $10,000 per year in continuing-education reimbursement.
- Employee Resource Groups and Inclusion Council participation.
- Flexible paid time off, 9 paid holidays, and a year-end week-long break.
- Paid parental leave: 20 weeks for primary caregivers and 14 weeks for secondary caregivers.
- Flexible return-to-work schedule.
- One-time $2,000 Calvin Care Cash incentive for eligible employees welcoming new family members.
- Flexible work environment with office or fully remote work options within the United States.

## Similar jobs

- [Staff+ Software Engineer, AI](https://hotfix.jobs/jobs/158dc955-461d-45ca-b5e2-61aa32ce5bca) - Aleph - Remote - $130k – $350k/yr
- [Staff AI Engineer](https://hotfix.jobs/jobs/ef3edbe2-3c9e-4733-8345-40c8d0307d13) - Grafana Labs - Remote - CA$164k – CA$197k/yr
- [Staff Applied Scientist, Personalization](https://hotfix.jobs/jobs/f3f16be4-af38-4045-b3a6-cb954a5272da) - OnePay - Remote - $180k – $250k/yr
- [Staff AI Enablement Engineer](https://hotfix.jobs/jobs/f66b54fc-c0f6-41bc-aa3e-dd03ad31f295) - Talkiatry - Remote - $190k – $230k/yr
- [Senior/Staff Engineer, Machine Learning - Online Mapping](https://hotfix.jobs/jobs/2dbc058c-d88d-4b51-8d5e-acacb5130b4c) - Nuro - Mountain View, CA - $194k – $291k/yr

**Apply:** https://hotfix.jobs/jobs/073cb8d3-c133-4565-a065-8eddbaf1c5b8
**Canonical:** https://hotfix.jobs/jobs/073cb8d3-c133-4565-a065-8eddbaf1c5b8