Skip to content
CohereCohere

Senior Research Scientist, Model Evaluation

Conducts research and builds benchmarks, methods, and infrastructure to evaluate the capabilities and progress of large language models. The role requires strong software engineering, prototyping, data-quality review, and rigorous measurement skills.

About the job

Responsibilities

  • Create ambitious new evaluation benchmarks that push the limits of what models can accomplish.
  • Work on highly cross-functional teams to translate model feedback into trustworthy, repeatable evaluations.
  • Conduct research to advance the state of the art in LLM evaluation methods, including:
    • Training LLM judges.
    • Refining LLM-based data synthesis pipelines.
    • Improving evaluation efficiency.
  • Build scalable and reusable tools for analyzing model performance.

Requirements and Qualifications

  • Experience rapidly building prototypes that demonstrate the boundaries of LLM capabilities.
  • Experience developing resources to measure model capabilities.
  • Extensive experience reviewing complex data and LLM outputs to ensure high data quality.
  • Strong focus on rigorously measuring AI capabilities and aligning measurements with the capabilities being evaluated.
  • Strong software engineering skills.

Compensation and Benefits

  • Weekly lunch stipend of $75, £75, or equivalent in local currency.
  • Full health and dental benefits, including a separate mental health budget.
  • RRSP matching, 401(k), or pension scheme.
  • 100% parental leave top-up for up to six months for either parent.
  • Annual enrichment benefits for arts and culture, fitness and wellness, quality time, and workspace improvements.
  • Education and learning stipend for conferences, courses, and coaching.
  • Six weeks of paid vacation.
  • Travel budget for remote employees visiting other offices and an annual company offsite.
  • Coworking benefit for employees who are not near an office.
  • $500 home office stipend.

Skills

LLMs, Llm Evaluation, Evaluation Benchmarks, Llm Judges, Data Synthesis, Software Engineering, Prototyping, Data Quality, Ai Capabilities, Model Performance

Decagon

Decagon

San Francisco, CA
Senior Research Engineer, Safety
$200k+/yrOn-site4+ YOEAI Research

Research and build safety models, evaluations, and runtime safeguards for conversational AI agents, addressing prompt injection, unsafe tool use, privacy, and policy risks. Requires 4+ years in AI/ML engineering, research, or safety plus experience deploying and evaluating language models or agentic systems.

Function Health

Function Health

United States

Senior Clinical Specialist, AI Systems
No salary listedRemoteAI Research

Evaluates and improves AI-generated clinical outputs, partnering with product and engineering teams to establish safety, accuracy, and clinical-quality standards. Requires an MD, DO, or equivalent clinical doctorate, substantial patient-care experience, strong clinical judgment, and the ability to learn AI evaluation techniques.

Gusto

Gusto

Denver, CO
AI Solutions Architect
$168k+/yrHybrid8+ YOEAI Research

Owns reusable patterns, standards, and tooling for production agentic service workflows, guiding platform priorities, automation measurement, and quality governance. Requires 8+ years in operations, product, or AI, hands-on agentic workflow experience, and strong LLM, metrics, and cross-functional influence skills.

Anthropic

Anthropic

San Francisco, CA
Applied AI, Research Engineer
$300k+/yrHybrid6+ YOEAI Research

Applied AI Research Engineer who tests model capabilities, builds demos and evaluations, supports strategic customer implementations, and translates field insights into product and research direction. Requires 6+ years of technical experience, programming proficiency, LLM development experience, and strong communication skills.

Ai2

Ai2

Seattle, WA

Senior Research Scientist, Open Ecosystem
$170k+/yrOn-site5+ YOEAI Research

Conduct research and build open foundation models and training systems aimed at accelerating scientific discovery. The role requires a PhD-level background and substantial experience training foundation models, with expertise in agentic training or multimodal data preferred.