Skip to content
IdmeIdme

Staff Software Engineer - AI Agent Evaluations

Leads the engineering discipline for evaluating, testing, and monitoring production AI agents, while building scalable eval infrastructure and developer tooling. Requires 8+ years of production software experience, strong backend skills, and expertise with LLM evaluation and agentic systems.

About the job

Responsibilities

  • Define standards for evaluating, validating, and monitoring AI agents, from prompt-based features to autonomous multi-step workflows.
  • Design and maintain evaluation pipelines for LLM outputs, agent behavior, tool use, and multi-turn interactions across development, staging, and production.
  • Instrument agentic systems for behavioral drift, regressions, and failure modes, including latency, correctness, hallucination rate, tool misuse, and policy adherence.
  • Lead testing strategies for non-deterministic AI systems, including red-teaming, golden datasets, LLM-as-judge pipelines, and property-based testing.
  • Build internal tooling, feedback loops, and testing workflows that improve the AI development inner loop and provide clear regression signals.
  • Establish engineering patterns, tooling, and education for responsible AI development and deployment.
  • Partner with Security, Platform, Product, and AI/ML teams to embed quality gates in agent workflows.
  • Mentor engineers on evaluation design, observability, and AI-specific testing.

Requirements

  • Bachelor's degree in Computer Science, Engineering, or equivalent experience.
  • 8+ years building and operating production software systems.
  • Experience evaluating or testing LLM-powered features or autonomous agents in production.
  • Proficiency with AI-assisted development tools such as Claude Code or Cursor.
  • Strong backend engineering fundamentals in Python, Java, Go, or equivalent.
  • Experience designing test infrastructure, CI/CD quality gates, or evaluation pipelines at scale.
  • Experience improving developer experience through internal tooling, toil reduction, or workflow acceleration.
  • Ability to lead cross-team technical initiatives and influence engineering standards.
  • Strong written and verbal communication across engineering, product, and leadership.
  • Experience building LLM-agent evaluation frameworks, including correctness graders, LLM-as-judge, human-in-the-loop evaluations, or benchmark dataset curation.
  • Familiarity with agentic frameworks such as Claude API, Anthropic SDK, BrainTrust, LangChain, LangGraph, CrewAI, or similar.
  • Production monitoring experience for AI systems, including behavioral drift detection, output sampling, or shadow scoring.
  • Red-teaming or adversarial testing experience for AI models or agents.

Nice-to-Haves

  • Experience in identity verification, fraud detection, or regulated industries.
  • Familiarity with Anthropic's model evaluation methodology or similar published evaluation research.
  • Experience applying Datadog or OpenTelemetry to AI workloads.
  • Track record of building developer tooling or platforms widely adopted by other teams.

Compensation and Benefits

  • Annual base salary: $217,565–$271,000 USD.
  • Comprehensive medical, dental, and vision coverage; health savings and flexible spending accounts; commuter benefits; life and disability insurance; 401(k) with company match; parental leave; paid time off and company holidays; and additional wellbeing, learning, childcare, pet insurance, and employee assistance benefits.

Skills

Python, Java, Go, RAG, Vector Search, Llm Evaluation, Llm-As-Judge, LangChain, LangGraph, Crewai, CI/CD, Datadog, OpenTelemetry, Red Teaming

Ironclad

Ironclad

San Francisco, CA

Senior Staff Software Engineer, Agentic Search
$220k+/yrHybrid10+ YOEML Engineering

Leads architecture and technical direction for agentic search systems combining LLMs, retrieval, and content-understanding pipelines for contract intelligence. The role requires 10+ years building production systems, deep search or LLM expertise, and strong cross-team technical leadership.

Shield AI

Shield AI

Washington, DC
Staff Engineer, Autonomy Capabilities – Maritime
$221k+/yrOn-site7+ YOEML Engineering

Leads development and integration of advanced maritime autonomy for USVs, UUVs, and cooperating UAVs, including motion planning, localization, safety, and multi-agent coordination. Requires staff-level technical leadership, substantial robotics experience, C++ and Python proficiency, and eligibility for a SECRET clearance.

Airbnb

Airbnb

United States

Staff Machine Learning Engineer, Relevance and Personalization
$212k+/yrRemote9+ YOEML Engineering

Staff machine learning engineer leading scalable ranking, search, recommendation, and personalization systems. The role requires 9+ years of applied machine learning experience, strong programming and data engineering skills, and expertise productionizing models and pipelines.

Airbnb

Airbnb

United States

Staff Machine Learning Engineer, Traffic Intelligence
$212k+/yrRemote9+ YOEML Engineering

Architects and operates production machine-learning systems that classify web and API traffic, detect bots and scrapers, and support real-time mitigation at internet edge latency. The role requires 9+ years of applied ML experience in adversarial domains and strong expertise in evaluation, data pipelines, and large-scale systems.

Harvey

Harvey

San Francisco, CA

Staff Software Engineer, Model Infrastructure
$231k+/yrHybrid7+ YOEML Engineering

Leads the design and operation of reliable, scalable model infrastructure powering AI inference across multiple providers. Requires 7+ years of distributed-systems engineering experience, strong programming skills, and expertise in production reliability and cloud infrastructure.