Skip to content
FetchFetch

Staff AI Evaluation Lead

Leads organization-wide AI evaluation and automation programs, establishing quality standards, metrics, production gates, and scalable evaluation infrastructure for LLM and agentic systems. The role requires 8+ years of relevant experience, strong technical systems expertise, and cross-functional leadership.

About the job

Responsibilities

  • Lead complex, high-impact automation and evaluation initiatives across workflows and teams.
  • Architect end-to-end workflows integrating datasets, evaluations, automations, and human-in-the-loop processes.
  • Build reusable harnesses, components, and evaluation pipelines that run against production without hands-on operation.
  • Define dataset standards, evaluation methodology, failure taxonomies, and quality measurement across AI Operations.
  • Create, maintain, and validate metric definitions against real production behavior; revise them as models, tooling, and architectures change.
  • Define production quality bars, determine where human review remains in the loop, and remove systems from production when they no longer meet standards.
  • Allocate evaluation depth according to system risk and make accepted tradeoffs visible to stakeholders.
  • Define when model or platform changes require organizational re-baselining and manage that process.
  • Evaluate new models and AI capabilities and determine adoption decisions.
  • Redesign workflows, tooling, and processes to improve performance and durability across AI Operations.
  • Partner with Engineering, Product, and AI teams to align priorities and deliver solutions.
  • Define success metrics, communicate performance and recommendations to leadership, and drive measurable business impact.
  • Mentor technical contributors and elevate the team’s automation and evaluation capabilities.
  • Stay current on emerging tools and approaches, piloting and translating them into actionable improvements.

Requirements

  • 8+ years of professional experience in AI, machine learning, operational automation, or a related field.
  • Proven ability to lead complex automation or evaluation initiatives across systems or teams.
  • Experience designing, building, and scaling evaluation frameworks, datasets, and quality systems for LLMs or AI products.
  • Fluency in SQL, JSON, APIs, and scripting with AI assistance, with ownership of the technical direction of an evaluation stack.
  • Deep working knowledge of LLM and agentic system behavior, automation platforms, and system design, demonstrated through systems built.
  • Experience with data pipelines, APIs, and production systems.
  • Experience influencing cross-functional stakeholders and aligning priorities.
  • Demonstrated ability to define metrics and drive measurable business impact.

Preferred Qualifications

  • Experience mentoring or leading technical contributors.
  • Experience evaluating agentic systems in production at scale.
  • Experience setting technical standards adopted across an organization.

Compensation and Benefits

  • Base salary range: $129,000 - $152,000.
  • Equity for full-time employees.
  • Dollar-for-dollar 401(k) match up to 4%.
  • Medical, dental, and vision plans, including pet coverage.
  • Up to $10,000 per year in continuing-education reimbursement.
  • Employee Resource Groups and Inclusion Council participation.
  • Flexible paid time off, 9 paid holidays, and a year-end week-long break.
  • Paid parental leave: 20 weeks for primary caregivers and 14 weeks for secondary caregivers.
  • Flexible return-to-work schedule.
  • One-time $2,000 Calvin Care Cash incentive for eligible employees welcoming new family members.
  • Flexible work environment with office or fully remote work options within the United States.

Skills

SQL, JSON, APIs, Python, LLMs, Agentic Systems, Machine Learning, Ai Evaluation, Data Pipelines, Automation Platforms, System Design, Production Systems

Aleph

Aleph

United States

Staff+ Software Engineer, AI
$130k+/yrRemote7+ YOEML Engineering

Own the shared AI foundation powering Aleph’s financial planning products, including model routing, context and tool systems, evaluations, and observability. The role requires Staff-level experience shipping production LLM and agentic systems, strong technical judgment, and a pragmatic builder’s mindset.

Grafana Labs

Grafana Labs

United States
Staff AI Engineer
CA$164k+/yrRemote8+ YOEML Engineering

Builds and owns production multi-agent AI infrastructure, backend integrations, and workflow automation for marketing operations. Requires 8+ years of software engineering experience, strong Python and JavaScript/Node.js skills, production LLM experience, and deep Google Cloud expertise.

OnePay

OnePay

United States

Staff Applied Scientist, Personalization
$180k+/yrRemote7+ YOEML Engineering

Design and productionize ML, deep learning, and LLM models for personalization, recommendations, and search systems. Requires 7+ years building production ML/AI with business impact and strong cross-functional collaboration skills.

Talkiatry

Talkiatry

United States

Staff AI Enablement Engineer
$190k+/yrRemote8+ YOEML Engineering

Staff-level engineer responsible for building AI agents and automation, evaluating developer AI tools, and driving adoption across the engineering organization. Requires 8+ years of software engineering experience plus production experience with LLMs, agentic systems, and applied machine learning.

Nuro

Nuro

Mountain View, CA

Senior/Staff Engineer, Machine Learning - Online Mapping
$194k+/yrOn-site7+ YOEML Engineering

Develop and productize online mapping models for autonomous navigation using real-world sensor data. The role requires deep ML expertise, robotics or computer vision experience, strong Python and deep learning framework skills, and a staff-level ability to deliver practical solutions.