Skip to content

Research Scientist – Frontier Evaluations

Designs, validates, and publishes rigorous evaluations and benchmarks for frontier AI systems across agentic, coding, safety, and expert-domain applications. The role requires strong research publication experience, scientific writing, experimental rigor, and the ability to deliver reproducible evaluation systems.

About the job

Responsibilities

  • Lead the end-to-end design, validation, launch, and continuous improvement of frontier AI benchmarks.
  • Partner with researchers and domain experts to develop evaluations focused on model failures, coverage gaps, and high-priority domains.
  • Analyze model capabilities and failure modes using rigorous experimental design and statistical methods.
  • Build reproducible evaluation systems, including harnesses, graders, and benchmark infrastructure.
  • Collaborate with researchers to post-train models and measure resulting performance gains.
  • Communicate results through benchmark reports, technical articles, and research papers.

Requirements

  • Strong record of publishing benchmarks or research papers.
  • Clear technical communication and strong scientific writing skills.
  • Commitment to experimental rigor, including baselines, ablations, statistical validity, and contamination controls.
  • Ability to take an ambiguous evaluation question from initial scoping through a reproducible public release.
  • Depth in agentic, coding, and safety evaluations, or applied machine learning in an expert domain.

Nice-to-haves

  • PhD in a related technical field.
  • Research publications at leading conferences or peer-reviewed journals.
  • Interest in multidisciplinary research and the creativity to combine methods and insights from AI, engineering, science, and other expert domains.

Compensation and Benefits

  • Salary: $210,000–$450,000 annually.
  • Meaningful equity.
  • Medical, vision, and dental insurance.
  • 401(k) with employer match.
  • Daily Uber Eats stipend.
  • Monthly wellness stipend.
  • Commuting costs covered.

Skills

Artificial Intelligence, Machine Learning, Benchmarking, Evaluation Harnesses, Statistical Methods, Experimental Design, Scientific Writing, Post-Training, Ablation Studies, Contamination Controls

Mercor

Mercor

San Francisco, CA

Research Scientist, APEX Benchmarks
$200k+/yrOn-siteAI Research

Leads the design, measurement, publication, and adoption of APEX benchmarks evaluating frontier models on economically valuable professional work. The role requires rigorous research judgment, strong coding and statistical skills, and excellent communication across technical, commercial, and research audiences.

Tessera Labs

Tessera Labs

San Jose, CA

Research Scientist
$200k+/yrOn-siteAI Research

Research Scientist defining and executing research on reliable long-horizon agents in enterprise environments. The role focuses on post-training and reinforcement learning, agent memory, evaluation, verification, and structured representations, combining hands-on experimentation with product delivery and publication.

Baseten

Baseten

San Francisco, CA

AI Engineer
$220k+/yrHybrid5+ YOEAI Research

Build and ship agentic AI product experiences, internal automation, and customer-facing features across the stack. The role requires 5+ years of software engineering experience, hands-on experience with AI or LLM-powered products, Python proficiency, and strong autonomy.

OpenAI

OpenAI

San Francisco, CA

People Research Scientist
$198k+/yrOn-siteAI Research

Conduct rigorous people research and applied data science to evaluate talent programs, organizational health, and employee experiences. The role requires advanced expertise in research design, experimentation, measurement, causal inference, statistical modeling, and responsible handling of sensitive employee data.

Snowflake

Snowflake

Bellevue, WA

AI Systems Research and Development Engineer – LLM Inference Systems & Optimization
$236k+/yrOn-site5+ YOEAI Research

Develop and optimize production-scale LLM inference systems across distributed runtimes, GPU kernels, scheduling, and model-system co-design. The role requires a bachelor’s degree and at least five years of experience in inference, distributed AI, GPU systems, or high-performance computing.