Skip to content
FetchFetchUnited States

Staff AI Evaluation Lead

Leads organization-wide AI evaluation and automation programs, establishing quality standards, metrics, production gates, and scalable evaluation infrastructure for LLM and agentic systems. The role requires 8+ years of relevant experience, strong technical systems expertise, and cross-functional leadership.

$129k – $152k/yr
Remote8+ YOEML Engineering

About the job

Responsibilities

  • Lead complex, high-impact automation and evaluation initiatives across workflows and teams.
  • Architect end-to-end workflows integrating datasets, evaluations, automations, and human-in-the-loop processes.
  • Build reusable harnesses, components, and evaluation pipelines that run against production without hands-on operation.
  • Define dataset standards, evaluation methodology, failure taxonomies, and quality measurement across AI Operations.
  • Create, maintain, and validate metric definitions against real production behavior; revise them as models, tooling, and architectures change.
  • Define production quality bars, determine where human review remains in the loop, and remove systems from production when they no longer meet standards.
  • Allocate evaluation depth according to system risk and make accepted tradeoffs visible to stakeholders.
  • Define when model or platform changes require organizational re-baselining and manage that process.
  • Evaluate new models and AI capabilities and determine adoption decisions.
  • Redesign workflows, tooling, and processes to improve performance and durability across AI Operations.
  • Partner with Engineering, Product, and AI teams to align priorities and deliver solutions.
  • Define success metrics, communicate performance and recommendations to leadership, and drive measurable business impact.
  • Mentor technical contributors and elevate the team’s automation and evaluation capabilities.
  • Stay current on emerging tools and approaches, piloting and translating them into actionable improvements.

Requirements

  • 8+ years of professional experience in AI, machine learning, operational automation, or a related field.
  • Proven ability to lead complex automation or evaluation initiatives across systems or teams.
  • Experience designing, building, and scaling evaluation frameworks, datasets, and quality systems for LLMs or AI products.
  • Fluency in SQL, JSON, APIs, and scripting with AI assistance, with ownership of the technical direction of an evaluation stack.
  • Deep working knowledge of LLM and agentic system behavior, automation platforms, and system design, demonstrated through systems built.
  • Experience with data pipelines, APIs, and production systems.
  • Experience influencing cross-functional stakeholders and aligning priorities.
  • Demonstrated ability to define metrics and drive measurable business impact.

Preferred Qualifications

  • Experience mentoring or leading technical contributors.
  • Experience evaluating agentic systems in production at scale.
  • Experience setting technical standards adopted across an organization.

Compensation and Benefits

  • Base salary range: $129,000 - $152,000.
  • Equity for full-time employees.
  • Dollar-for-dollar 401(k) match up to 4%.
  • Medical, dental, and vision plans, including pet coverage.
  • Up to $10,000 per year in continuing-education reimbursement.
  • Employee Resource Groups and Inclusion Council participation.
  • Flexible paid time off, 9 paid holidays, and a year-end week-long break.
  • Paid parental leave: 20 weeks for primary caregivers and 14 weeks for secondary caregivers.
  • Flexible return-to-work schedule.
  • One-time $2,000 Calvin Care Cash incentive for eligible employees welcoming new family members.
  • Flexible work environment with office or fully remote work options within the United States.

Skills

SQLJSONAPIsPythonLLMsAgentic SystemsMachine LearningAi EvaluationData PipelinesAutomation PlatformsSystem DesignProduction Systems