Skip to content
MercorMercor

Research Engineer – Benchmarking

Research Engineer focused on designing benchmarks, evaluation systems, rubrics, and failure-analysis workflows for frontier language models. The role requires strong applied AI research and coding experience, with expertise in model evaluation, data quality, and backend systems.

About the job

Responsibilities

  • Design, implement, and maintain benchmarks and metrics for tool use, agentic behavior, and real-world reasoning.
  • Build and operate end-to-end LLM evaluation systems, including runs, scoring, dashboards, and reporting.
  • Conduct systematic failure analysis of model outputs, categorize failure modes, quantify prevalence, and feed findings into reward design, data curation, and benchmark design.
  • Create and refine rubrics, automated evaluators, and scoring frameworks, balancing rigor with scalability across human evaluation and model-as-judge approaches.
  • Quantify data usability, quality, and impact on key benchmarks; use evaluations and failure analysis to guide data generation, augmentation, and curation.
  • Collaborate with AI researchers, applied AI teams, and data producers to align evaluations with training objectives.
  • Own benchmarks, evaluations, and failure-analysis workflows in a high-iteration research environment.

Requirements

  • Strong applied research background focused on model evaluation, benchmarking, and/or failure analysis.
  • Strong coding skills and hands-on experience with machine-learning models and evaluation code.
  • Solid understanding of data structures, algorithms, and backend systems.
  • Experience with APIs, SQL/NoSQL, and cloud platforms for running and storing evaluation results.
  • Ability to reason about model behavior, experimental results, and data quality.
  • Willingness to work in person in San Francisco five days per week.

Nice to Have

  • Industry experience on a post-training or evaluation/benchmarking team.
  • Publications at top-tier venues such as NeurIPS, ICML, or ACL, especially in evaluation or benchmarking.
  • Experience building or running LLM evaluations, benchmarks, or failure-analysis pipelines.
  • Experience with synthetic data generation, rubric design, or RL-style workflows using evaluations for reward shaping.
  • Work samples or code demonstrating relevant skills, such as evaluation frameworks, benchmark suites, failure-analysis reports, or tooling.

Compensation and Benefits

  • Bi-annual performance bonus structure.
  • Generous equity grant vested over four years.
  • Up to $15k relocation bonus.
  • $10K housing bonus for employees living within 0.5 miles of the office.
  • $1.5K monthly meal stipend.
  • Free Equinox membership.
  • $200 monthly laundry reimbursement.
  • $200 monthly personal wellness reimbursement.
  • Health, dental, and vision insurance.

Skills

Python, Machine Learning, Llm Evaluation, Benchmarking, Failure Analysis, Data Structures, Algorithms, Backend Systems, APIs, SQL, NoSQL, Cloud Platforms, Synthetic Data Generation, Reinforcement Learning, Neurips

AI Digest

AI Digest

Remote

Research Scientist - Member of Technical Staff
$150k+/yrRemoteAI Research

Conduct research on long-horizon, multi-agent AI behavior by designing agent environments, analyzing large-scale data, and running experiments. The role requires strong research judgment, rapid execution, independence, and familiarity with current AI developments.

AI Digest

AI Digest

Remote

Engineer - Member of Technical Staff
$150k+/yrRemoteAI Research

Build, optimize, and evaluate long-running and multi-agent AI systems, along with tools for monitoring and analyzing their real-world behavior. The role requires software engineering experience with coding agents, strong independence, and familiarity with current AI developments.

Counsel Health

Counsel Health

New York, NY
Research Scientist
$165k+/yrHybrid5+ YOEAI Research

Research Scientist developing and evaluating health-focused AI models, large language models, and agentic systems for clinical applications. The role requires advanced research experience, strong coding skills, healthcare or clinical-data experience, and top-tier AI/ML publications.

Hightouch

Hightouch

United States

Software Engineer, Applied AI Research
$180k+/yrRemote5+ YOEAI Research

Develops experimental AI techniques and prototypes for agentic marketing applications, with emphasis on image and video generation. The role requires strong backend or probabilistic systems expertise, quantitative thinking, creativity with LLM applications, and product intuition.

Greptile

Greptile

San Francisco, CA

Research Engineer
$180k+/yrOn-siteAI Research

The Research Engineer will apply advances in agents and language models to build and evaluate multi-agent systems for automated code validation and review. The role requires a computer science or equivalent background, research experience, strong programming skills, and product intuition.