Software Engineer, Agent Evaluation and Quality
Builds measurement, evaluation, and feedback infrastructure for AI agent quality. Designs datasets, pipelines, dashboards, and analysis tools to improve agent reliability, partnering with research and product teams. Requires strong data acumen and software engineering skills.
About the job
Responsibilities
- Design and build best-in-class AI evaluation system: curated datasets, offline replay, scorers / judges, regression alerts, and dashboards.
- Design feedback loops from real usage: collecting, cleaning, and interpreting user signals to inform model and harness changes.
- Develop analysis tooling and workflows for debugging agent behavior: deep dives on failure modes, clustering themes, and surfacing actionable insights.
- Improve reliability and guardrails by making quality measurable and operational: defining “good/bad/degraded” sessions, alerting, and triage primitives.
Requirements
- Built and operated evaluation or measurement systems, such as AI evals, experimentation, ranking/relevance, or search quality. Can turn ambiguous “quality” questions into concrete metrics, pipelines, and decisions.
- Strong data acumen, and can collaborate effectively with data scientists and researchers.
- Taste and strong opinions on model and agent behaviors. Stay up-to-date on emerging research and industry trends.
- Strong software engineering fundamentals and enjoy shipping production systems.
Skills
Ai Evaluation, Datasets, Scorers, Dashboards, Pipelines, Data Analysis, Machine Learning, Python, SQL, Experimentation
Similar jobs
Data Engineering jobsBuild and maintain dbt models, Snowflake semantic layers, and ingestion pipelines across business functions while improving data quality and resilience. The role requires 4–6 years of analytics or data engineering experience, strong dbt and SQL expertise, and a quantitative bachelor's degree.
Build and govern quote-to-cash data models and products integrating Salesforce, CPQ, billing, and finance systems. The role requires 5+ years of data engineering experience, strong SQL and Python skills, and expertise in self-service analytics for GTM teams.
Build and scale distributed data platforms, database systems, delivery services, and APIs, with emphasis on reliability, performance, observability, and data integrity. Requires 3+ years of software development experience with distributed systems and databases; Golang experience is preferred.
Oversee the lifecycle, quality, governance, and publication of research data across scientific programs. The role requires 3–5+ years of research data-management experience, strong metadata and FAIR-data expertise, and the ability to collaborate with researchers and engineers.
Own and evolve trusted data models for Marketing and Product use cases, from design and testing through monitoring and documentation. The role requires 3–5 years of data or analytics engineering experience, strong SQL and Python, dbt expertise, and Snowflake or comparable warehouse experience.