# Research Scientist, APEX Benchmarks

**Company:** [Mercor](https://hotfix.jobs/companies/mercor)
**Location:** San Francisco, CA
**Role:** AI Research
**Salary:** $200k – $400k/yr
**Skills:** Llm Evaluation, Benchmarking, Natural Language Processing, Experimental Design, Statistics, Python, Evaluation Harnesses, Contamination Controls, Inter-Rater Agreement, Model-As-Judge, Human Annotation, Data-Centric Ai, Machine Learning, Research Publishing
**Posted:** 2026-08-27

> Leads the design, measurement, publication, and adoption of APEX benchmarks evaluating frontier models on economically valuable professional work. The role requires rigorous research judgment, strong coding and statistical skills, and excellent communication across technical, commercial, and research audiences.

## Job Description

## Responsibilities
- Design the next generation of APEX benchmarks, including task taxonomies, difficulty calibration, contamination controls, and statistical design.
- Create expert-built datasets and grading rubrics at scale while maintaining defensible quality standards.
- Establish rigorous measurement practices, including confidence intervals, inter-rater agreement, human-versus-model-as-judge calibration, held-out splits, and failure analysis.
- Collaborate with academic and industry partners to co-design and promote benchmark adoption.
- Publish research through papers, open datasets, blog posts, conference talks, and leaderboard releases.
- Represent the company's research with frontier labs, customers, partners, the press, and the broader research community.
- Translate benchmark findings into technical reports, customer narratives, and go-to-market materials.
- Partner with data operations, engineering, product, and strategy to move benchmarks from design into production and inform the company roadmap.
- Track LLM evaluation research and incorporate relevant methods into benchmark development.

## Requirements
- Applied or academic research experience in LLM evaluation, benchmarking, NLP, or a related field.
- Track record of rigorous experimental design.
- Strong judgment about informative measurements of frontier-model behavior.
- Knowledge of sampling, variance, contamination, and grader reliability.
- Strong coding skills and ability to build evaluation harnesses, run experiments, and analyze results independently.
- Excellent written and verbal communication skills for technical and non-technical audiences.
- Comfort working in ambiguous, fast-moving, cross-functional environments.
- Interest in GTM strategy, startup dynamics, and the AI data business.
- Willingness to work onsite in the San Francisco office five days per week.

## Nice to Have
- Ph.D. in machine learning, NLP, or a related field, or equivalent industry or frontier-lab research experience.
- Publications at top-tier venues such as NeurIPS, ICML, ACL, or ICLR, particularly in evaluation, benchmarking, or data-centric AI.
- Experience authoring a widely adopted public benchmark or dataset.
- Industry experience on an evaluation, benchmarking, or post-training team at a frontier lab.
- Domain expertise in finance, law, consulting, accounting, medicine, or software engineering.
- Experience designing rubrics, model-as-judge pipelines, or large-scale human annotation programs.

## Compensation and Benefits
- Annual salary range: $200,000–$400,000.
- Bi-annual performance bonus structure.
- Generous equity grant vested over four years.
- Up to $15,000 relocation bonus.
- $10,000 housing bonus for employees living within 0.5 miles of the office.
- $1,500 monthly meal stipend.
- Equinox membership.
- $200 monthly laundry reimbursement.
- $200 monthly personal wellness reimbursement.
- Health, dental, and vision insurance.
- 401(k) with company match.

## Similar jobs

- [Research Scientist](https://hotfix.jobs/jobs/f5ff3391-4477-4eb8-b26f-78c72fc10328) - Tessera Labs - San Jose, CA - $200k – $300k/yr
- [People Research Scientist](https://hotfix.jobs/jobs/9e65a017-a0c0-4582-a774-59ce6f0c1e69) - OpenAI - San Francisco, CA - $198k – $220k/yr
- [Software Engineer (Gen AI)](https://hotfix.jobs/jobs/7af58525-5b44-47cc-a964-4a0eee5a3506) - Earnin - Mountain View, CA - $181k – $222k/yr
- [Machine Learning Research Scientist, Evaluations](https://hotfix.jobs/jobs/72d88855-3e74-43ec-82cd-9eaa9e95df69) - Scale AI - San Francisco, CA - $181k – $226k/yr
- [Machine Learning Research Scientist / Research Engineer, Post-Training](https://hotfix.jobs/jobs/aec3f24b-dccd-4095-b46f-79ea3c89be63) - Scale AI - San Francisco, CA - $181k – $226k/yr

**Apply:** https://hotfix.jobs/jobs/77bc636d-de61-47a6-9ca0-db586f4d2d22
**Canonical:** https://hotfix.jobs/jobs/77bc636d-de61-47a6-9ca0-db586f4d2d22