Skip to content

Research, Post-Training Evals

Researcher developing internal evaluations and research signals for AI post-training, including usability, correctness, auditing, agentic systems, and nuanced model behaviors. Requires evaluation experience, strong research judgment, Python, and familiarity with deep learning frameworks.

About the job

Responsibilities

  • Create internal evaluations and research signals for model capabilities and behaviors relevant to research and post-training.
  • Develop usability evaluations for real research and product workflows, and partner with data teams to improve training signals.
  • Improve evaluation correctness, including grader reliability, ambiguous ground truth, evaluator disagreement, false positives and negatives, and gaps between measured and intended behavior.
  • Build evaluation onboarding and auditing methodologies.
  • Develop evaluations that detect meaningful improvements, resist gaming, and generalize beyond their original benchmark or setup.
  • Develop specialized agentic evaluations, including harness development and cross-harness and cross-environment generalization studies.
  • Develop efficient, high-quality core-set signals for internal reinforcement-learning research.
  • Evaluate personalized preferences, biases, values, and other nuanced model behaviors.
  • Collaborate with post-training researchers, engineers, and the broader research organization.

Requirements

  • Bachelor's degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding.
  • Experience designing, building, or analyzing evaluations, benchmarks, datasets, graders, or other measurement systems.
  • Strong written and verbal communication skills.
  • Experience collaborating across research and engineering teams.
  • Python proficiency.
  • Familiarity with at least one deep learning framework, such as PyTorch, TensorFlow, or JAX.
  • Ability to debug distributed training and write scalable code.
  • Strong research judgment, including clean ablations, honest baselines, and clear technical writing.

Nice-to-haves

  • Experience with LLMs, post-training, reinforcement learning, or agentic systems.
  • Experience building AI evaluations, graders, benchmarks, or internal research signals.
  • Experience with evaluation correctness, auditing, human or LLM-based evaluation, or open-ended task evaluation.
  • Experience with agentic evaluation, harnesses, long-horizon tasks, or cross-environment generalization.
  • Experience evaluating preferences, personalization, biases, values, or other nuanced model behaviors.
  • Track record of developing evaluation methodologies or research signals that influenced model development.
  • PhD in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline, or equivalent industry research experience.

Skills

Python, PyTorch, TensorFlow, JAX, LLMs, Reinforcement Learning, Agentic Systems, Evaluation Design, Benchmarking, Distributed Training

Improbable

Improbable

Remote

AI Researcher
No salary listedRemoteAI Research

Conduct applied research on AI agents, designing experiments and evaluation systems to improve reliability, context retention, and multi-step task completion. The role requires strong AI/ML research, engineering, experimental design, and communication skills.

Anthropic

Anthropic

San Francisco, CA

Research Engineer, Takeoff Intel
$350k+/yrHybridAI Research

Research Engineer building large-scale AI capability evaluations, telemetry, data pipelines, and analysis tools for Anthropic’s Takeoff Intel team. The role requires hands-on large language model experimentation, rapid prototyping, data expertise, and strong research collaboration.

Sardine

Sardine

United States

Applied AI Research Scientist
No salary listedRemote4+ YOEAI Research

Conduct applied research on foundation models for fraud detection using large-scale behavioral and financial-risk data. The role spans experimentation, evaluation, production deployment, and cross-functional work on model governance, requiring 4+ years of applied ML experience and strong Python and SQL skills.

OpenAI

OpenAI

San Francisco, CA

Researcher, Agent Safety, Oversight and System Mitigations
$380k+/yrHybridAI Research

Researcher or engineer focused on designing, evaluating, and productionizing oversight systems and safety mitigations for autonomous AI agents. The role requires strong systems or security reasoning, threat-modeling ability, and experience building practical evaluations and controls.

OpenAI

OpenAI

San Francisco, CA

Researcher, Agent Safety, Training and Evaluations
$380k+/yrHybridAI Research

Researcher focused on training and evaluating frontier AI agents, mining incidents, and building scalable safety measurement systems. The role requires strong research or ML engineering execution, quantitative judgment, and the ability to own ambiguous projects end to end.