Skip to content
Distyl AIDistyl AI

Applied AI Researcher, Benchmarking

Designs and constructs AI benchmarks and evaluation frameworks to measure reasoning, reliability, and real-world impact of intelligent systems. Requires experience with model evaluations, statistical rigor, building with AI models, and strong programming for prototypes.

About the job

Key Responsibilities

  • Design evaluation frameworks that capture reasoning depth, interaction quality, reliability, and operational impact.
  • Construct benchmarks that reflect real-world complexity to judge new architectures, techniques, and releases.
  • Explore new paradigms for evaluating intelligent systems: adversarial robustness testing, longitudinal performance tracking, and human-in-the-loop assessment.
  • Investigate how metrics shape model behavior and establish rigorous methodologies for quantifying emergent capability.

Who You Are (Requirements)

  • Experience designing and running evaluations: built or maintained benchmarks, test suites, or experimental frameworks.
  • Statistical and analytical rigor: design fair, reproducible experiments and extract signal from noisy results.
  • Experience building with models (compound AI systems, agentic collaboration, ensembling, ReAct, graph-of-thoughts, etc.).
  • Proven track record of research results (publications, public work).
  • Uses AI every day (ChatGPT, Cursor, Perplexity).
  • Strong programming and data analysis skills for prototypes and experiments.
  • Biases towards showing vs telling.

Compensation & Benefits

  • Base salary: $150K – $250K (depending on experience, location, level).
  • Meaningful equity.
  • 100% covered medical, dental, vision for employees/dependents.
  • 401(k), commuter benefits, in-office lunch.
  • Access to state-of-the-art models and AI tools.

Skills

Ai Benchmarks, Evaluation Frameworks, Llm Evaluation, React, Graph-Of-Thoughts, Ensembling, Python, Data Analysis, Compound Ai Systems, Agentic Systems

AI Digest

AI Digest

Remote

Research Scientist - Member of Technical Staff
$150k+/yrRemoteAI Research

Conduct research on long-horizon, multi-agent AI behavior by designing agent environments, analyzing large-scale data, and running experiments. The role requires strong research judgment, rapid execution, independence, and familiarity with current AI developments.

AI Digest

AI Digest

Remote

Engineer - Member of Technical Staff
$150k+/yrRemoteAI Research

Build, optimize, and evaluate long-running and multi-agent AI systems, along with tools for monitoring and analyzing their real-world behavior. The role requires software engineering experience with coding agents, strong independence, and familiarity with current AI developments.

Counsel Health

Counsel Health

New York, NY
Research Scientist
$165k+/yrHybrid5+ YOEAI Research

Research Scientist developing and evaluating health-focused AI models, large language models, and agentic systems for clinical applications. The role requires advanced research experience, strong coding skills, healthcare or clinical-data experience, and top-tier AI/ML publications.

Mercor

Mercor

San Francisco, CA

Research Engineer – Benchmarking
$130k+/yrOn-siteAI Research

Research Engineer focused on designing benchmarks, evaluation systems, rubrics, and failure-analysis workflows for frontier language models. The role requires strong applied AI research and coding experience, with expertise in model evaluation, data quality, and backend systems.

Hightouch

Hightouch

United States

Software Engineer, Applied AI Research
$180k+/yrRemote5+ YOEAI Research

Develops experimental AI techniques and prototypes for agentic marketing applications, with emphasis on image and video generation. The role requires strong backend or probabilistic systems expertise, quantitative thinking, creativity with LLM applications, and product intuition.