Skip to content
Scale AIScale AI

Machine Learning Research Scientist, Evaluations

Research Scientist focused on evaluating frontier language and multimodal models, diagnosing failure modes, and building rigorous benchmarks. The role requires advanced training in AI or a related field, post-training expertise, and published machine learning research.

About the job

Responsibilities

  • Analyze model behavior to identify, characterize, and diagnose failure modes in frontier large language models and agents, including capability gaps, reasoning errors, robustness issues, and alignment issues.
  • Design and build benchmarks and evaluation methods for LLM capabilities across text and multimodal modalities.
  • Apply post-training expertise to connect observed failures with data and training interventions.
  • Collaborate with researchers and engineers to define evaluation-driven AI development best practices.
  • Partner with foundation model labs to translate failure analyses into technical and strategic input for next-generation generative AI models.
  • Publish research findings in top-tier AI conferences.

Requirements

  • Ph.D. or master's degree in computer science, machine learning, AI, or a related field.
  • Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning.
  • Experience with post-training techniques such as supervised fine-tuning, RLHF, preference modeling, or instruction tuning.
  • Experience with LLM evaluation or benchmark development.
  • Published machine learning research in major conferences or journals.
  • Excellent written and verbal communication skills.
  • Previous experience in a customer-facing role.

Compensation

  • Base salary range: $180,600–$225,750 USD.
  • Compensation may include equity and benefits, including health, dental and vision coverage, retirement benefits, learning and development stipend, PTO, and potentially a commuter stipend.

Skills

LLMs, Supervised Fine-Tuning, RLHF, Reward Modeling, Preference Modeling, Instruction Tuning, Deep Learning, Reinforcement Learning, Model Fine-Tuning, Llm Evaluation, Benchmark Development, Multimodal Models, Python, Failure Analysis, Research Publications

Scale AI

Scale AI

San Francisco, CA
Machine Learning Research Scientist / Research Engineer, Post-Training
$181k+/yrOn-siteAI Research

Research novel post-training methods for large language models, focusing on preference optimization, data curation, evaluation, alignment, and robustness across text and multimodal systems. Requires advanced academic training and experience with deep learning, reinforcement learning, and post-training techniques.

Earnin

Earnin

Mountain View, CA

Software Engineer (Gen AI)
$181k+/yrHybrid3+ YOEAI Research

Build agent-driven chatbots and generative AI workflows for financial-wellness products, owning features from design through impact measurement. The role requires at least three years of software engineering experience, strong system design, maintainable coding practices, and a bachelor’s degree or equivalent experience.

Hightouch

Hightouch

United States

Software Engineer, Applied AI Research
$180k+/yrRemote5+ YOEAI Research

Develops experimental AI techniques and prototypes for agentic marketing applications, with emphasis on image and video generation. The role requires strong backend or probabilistic systems expertise, quantitative thinking, creativity with LLM applications, and product intuition.

Greptile

Greptile

San Francisco, CA

Research Engineer
$180k+/yrOn-siteAI Research

The Research Engineer will apply advances in agents and language models to build and evaluate multi-agent systems for automated code validation and review. The role requires a computer science or equivalent background, research experience, strong programming skills, and product intuition.

Counsel Health

Counsel Health

New York, NY
Research Scientist
$165k+/yrHybrid5+ YOEAI Research

Research Scientist developing and evaluating health-focused AI models, large language models, and agentic systems for clinical applications. The role requires advanced research experience, strong coding skills, healthcare or clinical-data experience, and top-tier AI/ML publications.