# Machine Learning Research Scientist, Evaluations

**Company:** [Scale AI](https://hotfix.jobs/companies/scale-ai)
**Location:** San Francisco, CA, Seattle, WA, New York, NY
**Role:** AI Research
**Salary:** $181k – $226k/yr
**Skills:** LLMs, Supervised Fine-Tuning, RLHF, Reward Modeling, Preference Modeling, Instruction Tuning, Deep Learning, Reinforcement Learning, Model Fine-Tuning, Llm Evaluation, Benchmark Development, Multimodal Models, Python, Failure Analysis, Research Publications
**Posted:** 2026-08-26

> Research Scientist focused on evaluating frontier language and multimodal models, diagnosing failure modes, and building rigorous benchmarks. The role requires advanced training in AI or a related field, post-training expertise, and published machine learning research.

## Job Description

## Responsibilities
- Analyze model behavior to identify, characterize, and diagnose failure modes in frontier large language models and agents, including capability gaps, reasoning errors, robustness issues, and alignment issues.
- Design and build benchmarks and evaluation methods for LLM capabilities across text and multimodal modalities.
- Apply post-training expertise to connect observed failures with data and training interventions.
- Collaborate with researchers and engineers to define evaluation-driven AI development best practices.
- Partner with foundation model labs to translate failure analyses into technical and strategic input for next-generation generative AI models.
- Publish research findings in top-tier AI conferences.

## Requirements
- Ph.D. or master's degree in computer science, machine learning, AI, or a related field.
- Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning.
- Experience with post-training techniques such as supervised fine-tuning, RLHF, preference modeling, or instruction tuning.
- Experience with LLM evaluation or benchmark development.
- Published machine learning research in major conferences or journals.
- Excellent written and verbal communication skills.
- Previous experience in a customer-facing role.

## Compensation
- Base salary range: **$180,600–$225,750 USD**.
- Compensation may include equity and benefits, including health, dental and vision coverage, retirement benefits, learning and development stipend, PTO, and potentially a commuter stipend.

## Similar jobs

- [Machine Learning Research Scientist / Research Engineer, Post-Training](https://hotfix.jobs/jobs/aec3f24b-dccd-4095-b46f-79ea3c89be63) - Scale AI - San Francisco, CA - $181k – $226k/yr
- [Software Engineer (Gen AI)](https://hotfix.jobs/jobs/7af58525-5b44-47cc-a964-4a0eee5a3506) - Earnin - Mountain View, CA - $181k – $222k/yr
- [Software Engineer, Applied AI Research](https://hotfix.jobs/jobs/b270e535-825d-433e-a8e9-ea9a8e8d2836) - Hightouch - Remote - $180k – $400k/yr
- [Research Engineer](https://hotfix.jobs/jobs/996c82c8-13bd-4ba2-91e1-a50a6499eed0) - Greptile - San Francisco, CA - $180k – $280k/yr
- [Research Scientist](https://hotfix.jobs/jobs/f8da4cc2-a216-4a79-b247-5d2bb83ac27e) - Counsel Health - New York, NY - $165k – $220k/yr

**Apply:** https://hotfix.jobs/jobs/72d88855-3e74-43ec-82cd-9eaa9e95df69
**Canonical:** https://hotfix.jobs/jobs/72d88855-3e74-43ec-82cd-9eaa9e95df69