# Research, Post-Training Evals

**Company:** [Thinking Machines Lab](https://hotfix.jobs/companies/thinking-machines-lab)
**Location:** San Francisco, CA
**Role:** AI Research
**Skills:** Python, PyTorch, TensorFlow, JAX, LLMs, Reinforcement Learning, Agentic Systems, Evaluation Design, Benchmarking, Distributed Training
**Posted:** 2026-08-27

> Researcher developing internal evaluations and research signals for AI post-training, including usability, correctness, auditing, agentic systems, and nuanced model behaviors. Requires evaluation experience, strong research judgment, Python, and familiarity with deep learning frameworks.

## Job Description

## Responsibilities
- Create internal evaluations and research signals for model capabilities and behaviors relevant to research and post-training.
- Develop usability evaluations for real research and product workflows, and partner with data teams to improve training signals.
- Improve evaluation correctness, including grader reliability, ambiguous ground truth, evaluator disagreement, false positives and negatives, and gaps between measured and intended behavior.
- Build evaluation onboarding and auditing methodologies.
- Develop evaluations that detect meaningful improvements, resist gaming, and generalize beyond their original benchmark or setup.
- Develop specialized agentic evaluations, including harness development and cross-harness and cross-environment generalization studies.
- Develop efficient, high-quality core-set signals for internal reinforcement-learning research.
- Evaluate personalized preferences, biases, values, and other nuanced model behaviors.
- Collaborate with post-training researchers, engineers, and the broader research organization.

## Requirements
- Bachelor's degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding.
- Experience designing, building, or analyzing evaluations, benchmarks, datasets, graders, or other measurement systems.
- Strong written and verbal communication skills.
- Experience collaborating across research and engineering teams.
- Python proficiency.
- Familiarity with at least one deep learning framework, such as PyTorch, TensorFlow, or JAX.
- Ability to debug distributed training and write scalable code.
- Strong research judgment, including clean ablations, honest baselines, and clear technical writing.

## Nice-to-haves
- Experience with LLMs, post-training, reinforcement learning, or agentic systems.
- Experience building AI evaluations, graders, benchmarks, or internal research signals.
- Experience with evaluation correctness, auditing, human or LLM-based evaluation, or open-ended task evaluation.
- Experience with agentic evaluation, harnesses, long-horizon tasks, or cross-environment generalization.
- Experience evaluating preferences, personalization, biases, values, or other nuanced model behaviors.
- Track record of developing evaluation methodologies or research signals that influenced model development.
- PhD in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline, or equivalent industry research experience.

## Similar jobs

- [AI Researcher](https://hotfix.jobs/jobs/077574a5-3c6d-4634-a23c-4a909ce8aa65) - Improbable - Remote
- [Research Engineer, Takeoff Intel](https://hotfix.jobs/jobs/398824a2-65cc-4e28-aaeb-26c3b6610876) - Anthropic - San Francisco, CA - $350k – $850k/yr
- [Applied AI Research Scientist](https://hotfix.jobs/jobs/93baef6f-91a8-4c62-acaa-44c3ea48b467) - Sardine - Remote
- [Researcher, Agent Safety, Oversight and System Mitigations](https://hotfix.jobs/jobs/4544e3bb-bb96-43d2-a96c-cd364b641660) - OpenAI - San Francisco, CA - $380k – $500k/yr
- [Researcher, Agent Safety, Training and Evaluations](https://hotfix.jobs/jobs/d80336da-e453-4999-9f26-85a125b679d9) - OpenAI - San Francisco, CA - $380k – $500k/yr

**Apply:** https://hotfix.jobs/jobs/c1d5c6e7-eaf9-4258-b839-f2a29c1a105c
**Canonical:** https://hotfix.jobs/jobs/c1d5c6e7-eaf9-4258-b839-f2a29c1a105c