# Research Engineer – Benchmarking

**Company:** [Mercor](https://hotfix.jobs/companies/mercor)
**Location:** San Francisco, CA
**Role:** AI Research
**Salary:** $130k – $500k/yr
**Skills:** Python, Machine Learning, Llm Evaluation, Benchmarking, Failure Analysis, Data Structures, Algorithms, Backend Systems, APIs, SQL, NoSQL, Cloud Platforms, Synthetic Data Generation, Reinforcement Learning, Neurips
**Posted:** 2026-08-18

> Research Engineer focused on designing benchmarks, evaluation systems, rubrics, and failure-analysis workflows for frontier language models. The role requires strong applied AI research and coding experience, with expertise in model evaluation, data quality, and backend systems.

## Job Description

## Responsibilities

- Design, implement, and maintain benchmarks and metrics for tool use, agentic behavior, and real-world reasoning.
- Build and operate end-to-end LLM evaluation systems, including runs, scoring, dashboards, and reporting.
- Conduct systematic failure analysis of model outputs, categorize failure modes, quantify prevalence, and feed findings into reward design, data curation, and benchmark design.
- Create and refine rubrics, automated evaluators, and scoring frameworks, balancing rigor with scalability across human evaluation and model-as-judge approaches.
- Quantify data usability, quality, and impact on key benchmarks; use evaluations and failure analysis to guide data generation, augmentation, and curation.
- Collaborate with AI researchers, applied AI teams, and data producers to align evaluations with training objectives.
- Own benchmarks, evaluations, and failure-analysis workflows in a high-iteration research environment.

## Requirements

- Strong applied research background focused on model evaluation, benchmarking, and/or failure analysis.
- Strong coding skills and hands-on experience with machine-learning models and evaluation code.
- Solid understanding of data structures, algorithms, and backend systems.
- Experience with APIs, SQL/NoSQL, and cloud platforms for running and storing evaluation results.
- Ability to reason about model behavior, experimental results, and data quality.
- Willingness to work in person in San Francisco five days per week.

## Nice to Have

- Industry experience on a post-training or evaluation/benchmarking team.
- Publications at top-tier venues such as NeurIPS, ICML, or ACL, especially in evaluation or benchmarking.
- Experience building or running LLM evaluations, benchmarks, or failure-analysis pipelines.
- Experience with synthetic data generation, rubric design, or RL-style workflows using evaluations for reward shaping.
- Work samples or code demonstrating relevant skills, such as evaluation frameworks, benchmark suites, failure-analysis reports, or tooling.

## Compensation and Benefits

- Bi-annual performance bonus structure.
- Generous equity grant vested over four years.
- Up to $15k relocation bonus.
- $10K housing bonus for employees living within 0.5 miles of the office.
- $1.5K monthly meal stipend.
- Free Equinox membership.
- $200 monthly laundry reimbursement.
- $200 monthly personal wellness reimbursement.
- Health, dental, and vision insurance.

## Similar jobs

- [Research Scientist - Member of Technical Staff](https://hotfix.jobs/jobs/9807d0a2-a252-48f0-81fa-2175608c94b8) - AI Digest - Remote - $150k – $350k/yr
- [Engineer - Member of Technical Staff](https://hotfix.jobs/jobs/cba82dd2-84d4-48a7-8d75-58783156ab95) - AI Digest - Remote - $150k – $350k/yr
- [Research Scientist](https://hotfix.jobs/jobs/f8da4cc2-a216-4a79-b247-5d2bb83ac27e) - Counsel Health - New York, NY - $165k – $220k/yr
- [Software Engineer, Applied AI Research](https://hotfix.jobs/jobs/b270e535-825d-433e-a8e9-ea9a8e8d2836) - Hightouch - Remote - $180k – $400k/yr
- [Research Engineer](https://hotfix.jobs/jobs/996c82c8-13bd-4ba2-91e1-a50a6499eed0) - Greptile - San Francisco, CA - $180k – $280k/yr

**Apply:** https://hotfix.jobs/jobs/92ac70e0-232e-4223-b628-2416fc91ef52
**Canonical:** https://hotfix.jobs/jobs/92ac70e0-232e-4223-b628-2416fc91ef52