Staff Software Engineer - AI Agent Evaluations
Leads the engineering discipline for evaluating, testing, and monitoring production AI agents, while building scalable eval infrastructure and developer tooling. Requires 8+ years of production software experience, strong backend skills, and expertise with LLM evaluation and agentic systems.
About the job
Responsibilities
- Define standards for evaluating, validating, and monitoring AI agents, from prompt-based features to autonomous multi-step workflows.
- Design and maintain evaluation pipelines for LLM outputs, agent behavior, tool use, and multi-turn interactions across development, staging, and production.
- Instrument agentic systems for behavioral drift, regressions, and failure modes, including latency, correctness, hallucination rate, tool misuse, and policy adherence.
- Lead testing strategies for non-deterministic AI systems, including red-teaming, golden datasets, LLM-as-judge pipelines, and property-based testing.
- Build internal tooling, feedback loops, and testing workflows that improve the AI development inner loop and provide clear regression signals.
- Establish engineering patterns, tooling, and education for responsible AI development and deployment.
- Partner with Security, Platform, Product, and AI/ML teams to embed quality gates in agent workflows.
- Mentor engineers on evaluation design, observability, and AI-specific testing.
Requirements
- Bachelor's degree in Computer Science, Engineering, or equivalent experience.
- 8+ years building and operating production software systems.
- Experience evaluating or testing LLM-powered features or autonomous agents in production.
- Proficiency with AI-assisted development tools such as Claude Code or Cursor.
- Strong backend engineering fundamentals in Python, Java, Go, or equivalent.
- Experience designing test infrastructure, CI/CD quality gates, or evaluation pipelines at scale.
- Experience improving developer experience through internal tooling, toil reduction, or workflow acceleration.
- Ability to lead cross-team technical initiatives and influence engineering standards.
- Strong written and verbal communication across engineering, product, and leadership.
- Experience building LLM-agent evaluation frameworks, including correctness graders, LLM-as-judge, human-in-the-loop evaluations, or benchmark dataset curation.
- Familiarity with agentic frameworks such as Claude API, Anthropic SDK, BrainTrust, LangChain, LangGraph, CrewAI, or similar.
- Production monitoring experience for AI systems, including behavioral drift detection, output sampling, or shadow scoring.
- Red-teaming or adversarial testing experience for AI models or agents.
Nice-to-Haves
- Experience in identity verification, fraud detection, or regulated industries.
- Familiarity with Anthropic's model evaluation methodology or similar published evaluation research.
- Experience applying Datadog or OpenTelemetry to AI workloads.
- Track record of building developer tooling or platforms widely adopted by other teams.
Compensation and Benefits
- Annual base salary: $217,565–$271,000 USD.
- Comprehensive medical, dental, and vision coverage; health savings and flexible spending accounts; commuter benefits; life and disability insurance; 401(k) with company match; parental leave; paid time off and company holidays; and additional wellbeing, learning, childcare, pet insurance, and employee assistance benefits.
Skills
Python, Java, Go, RAG, Vector Search, Llm Evaluation, Llm-As-Judge, LangChain, LangGraph, Crewai, CI/CD, Datadog, OpenTelemetry, Red Teaming
Similar jobs
ML Engineering jobsLeads architecture and technical direction for agentic search systems combining LLMs, retrieval, and content-understanding pipelines for contract intelligence. The role requires 10+ years building production systems, deep search or LLM expertise, and strong cross-team technical leadership.
Leads development and integration of advanced maritime autonomy for USVs, UUVs, and cooperating UAVs, including motion planning, localization, safety, and multi-agent coordination. Requires staff-level technical leadership, substantial robotics experience, C++ and Python proficiency, and eligibility for a SECRET clearance.
Staff machine learning engineer leading scalable ranking, search, recommendation, and personalization systems. The role requires 9+ years of applied machine learning experience, strong programming and data engineering skills, and expertise productionizing models and pipelines.
Architects and operates production machine-learning systems that classify web and API traffic, detect bots and scrapers, and support real-time mitigation at internet edge latency. The role requires 9+ years of applied ML experience in adversarial domains and strong expertise in evaluation, data pipelines, and large-scale systems.
Leads the design and operation of reliable, scalable model infrastructure powering AI inference across multiple providers. Requires 7+ years of distributed-systems engineering experience, strong programming skills, and expertise in production reliability and cloud infrastructure.