Research Engineer, Benchmarks
Build high-quality, domain-specific benchmarks and infrastructure to rigorously evaluate frontier AI agents on realistic workflows. Requires strong Python/Docker/Linux skills, experience with evals or benchmarks, and a deep understanding of what makes a benchmark reliable and useful.
About the job
Responsibilities
- Own the design, implementation, and quality of HUD’s internal agent benchmarks
- Work with subject-matter experts to define tasks and create domain-specific benchmarks that evaluate agents on realistic workflows
- Build infrastructure to reliably run models and agents against benchmark tasks
- Develop metrics and analyses to understand benchmark difficulty, reliability, and failure modes
- Validate whether benchmark performance correlates with real-world evals, customer needs, and lab expectations
- Write clear documentation and benchmark reports that make results legible and credible to technical audiences
Requirements
- Proficiency in Python, Docker, and Linux environments
- Published papers or written technical blogs on relevant topics such as public benchmarks and their limitations, model failure modes, etc.
- Strong understanding of what a “good benchmark” means and what makes one realistic, reliable, and useful
- Experience working on environments and evals
- Curiosity and ability to truly understand how workflows in various domains work
Nice-to-Haves
- Detail-oriented and able to spot subtle inconsistencies or edge cases in tasks
- Able to reason from first principles about task design, scoring, and failure modes
- Thrive in unstructured problem spaces
- Early-stage startup experience with ability to work independently in fast-paced environments
- Strong communication skills for remote collaboration across time zones
Skills
Python, Docker, Linux, Benchmarks, Evals, Agent Evaluation
Similar jobs
AI Research jobsConduct applied research on AI agents, designing experiments and evaluation systems to improve reliability, context retention, and multi-step task completion. The role requires strong AI/ML research, engineering, experimental design, and communication skills.
Research Engineer building large-scale AI capability evaluations, telemetry, data pipelines, and analysis tools for Anthropic’s Takeoff Intel team. The role requires hands-on large language model experimentation, rapid prototyping, data expertise, and strong research collaboration.
Conduct applied research on foundation models for fraud detection using large-scale behavioral and financial-risk data. The role spans experimentation, evaluation, production deployment, and cross-functional work on model governance, requiring 4+ years of applied ML experience and strong Python and SQL skills.
Researcher or engineer focused on designing, evaluating, and productionizing oversight systems and safety mitigations for autonomous AI agents. The role requires strong systems or security reasoning, threat-modeling ability, and experience building practical evaluations and controls.
Researcher focused on training and evaluating frontier AI agents, mining incidents, and building scalable safety measurement systems. The role requires strong research or ML engineering execution, quantitative judgment, and the ability to own ambiguous projects end to end.