Skip to content
MercorMercor

Mercor Research Fellowship - APEX

Research fellows propose, build, validate, and publish benchmarks or evaluation methodologies for measuring frontier AI performance on economically valuable professional and scientific work. The fellowship requires a specific research pitch, relevant technical or adjacent-field background, and a commitment of at least 20 hours per week.

About the job

What You’ll Do

  • Propose and scope a new benchmark or evaluation technique in an under-covered domain, or a meaningfully harder version of an existing APEX benchmark.
  • Design task specifications and grading rubrics with vetted domain experts, including lawyers, accountants, engineers, scientists, and consultants.
  • Build and validate benchmarks by piloting tasks, calibrating scoring, and stress-testing for contamination and gameable shortcuts.
  • Run frontier models against benchmarks and analyze where and why they fail.
  • Publish results as a paper, open dataset, APEX leaderboard, or methodology adopted internally.
  • Partner with research and engineering teams to incorporate findings into APEX’s public benchmark family.

Focus Areas

  • Long-horizon, multi-application agentic tasks in professional services, including law, finance, and consulting.
  • Real-world software engineering evaluation beyond issue resolution.
  • Professional accounting and finance workflows.
  • AI-for-science evaluations in mathematics, biology, materials science, and theoretical physics.
  • Evaluation methodology, including contamination resistance, rubric design, human-versus-model grading agreement, and cost-adjusted scoring.

Requirements

  • Genuine interest in evaluation as a research discipline.
  • Background in computer science, machine learning, statistics, measurement, psychometrics, HCI, social science, or an adjacent field.
  • A specific, well-scoped benchmark or evaluation-technique idea.
  • Comfort working in a fast-paced startup environment with limited hand-holding.
  • Ability to commit at least 20 hours per week for the fellowship duration.

Nice-to-Haves

  • Experience with agentic evaluation or reinforcement-learning environments.
  • Domain expertise in law, finance, medicine, or a scientific field.

Compensation & Benefits

  • $40,000 stipend for 3 months or $80,000 stipend for 6 months.
  • Unlimited API credits.
  • Dedicated budget for GPU compute and paid expert or human-data time.
  • Weekly one-on-one mentorship with an APEX research team member and regular access to the broader research organization.
  • Access to frontier model APIs and Mercor’s internal evaluation infrastructure.
  • Optional desk at Mercor’s San Francisco office.
  • Introductions to researchers across frontier labs and academia.
  • Standout fellows may be considered for a full-time offer on the APEX research team.

Skills

Artificial Intelligence, Machine Learning, Statistics, Benchmarking, Evaluation Methodology, Agentic Evaluation, Reinforcement Learning, Rubric Design, Python, Software Engineering, Accounting, Psychometrics, Hci, Gpu Computing, Scientific Computing

Mercor

Mercor

San Francisco, CA

Research Engineer – Benchmarking
$130k+/yrOn-siteAI Research

Research Engineer focused on designing benchmarks, evaluation systems, rubrics, and failure-analysis workflows for frontier language models. The role requires strong applied AI research and coding experience, with expertise in model evaluation, data quality, and backend systems.

AI Digest

AI Digest

Remote

Research Scientist - Member of Technical Staff
$150k+/yrRemoteAI Research

Conduct research on long-horizon, multi-agent AI behavior by designing agent environments, analyzing large-scale data, and running experiments. The role requires strong research judgment, rapid execution, independence, and familiarity with current AI developments.

AI Digest

AI Digest

Remote

Engineer - Member of Technical Staff
$150k+/yrRemoteAI Research

Build, optimize, and evaluate long-running and multi-agent AI systems, along with tools for monitoring and analyzing their real-world behavior. The role requires software engineering experience with coding agents, strong independence, and familiarity with current AI developments.

Counsel Health

Counsel Health

New York, NY
Research Scientist
$165k+/yrHybrid5+ YOEAI Research

Research Scientist developing and evaluating health-focused AI models, large language models, and agentic systems for clinical applications. The role requires advanced research experience, strong coding skills, healthcare or clinical-data experience, and top-tier AI/ML publications.

Hightouch

Hightouch

United States

Software Engineer, Applied AI Research
$180k+/yrRemote5+ YOEAI Research

Develops experimental AI techniques and prototypes for agentic marketing applications, with emphasis on image and video generation. The role requires strong backend or probabilistic systems expertise, quantitative thinking, creativity with LLM applications, and product intuition.