# Research, Safety

**Company:** [Thinking Machines Lab](https://hotfix.jobs/companies/thinking-machines-lab)
**Location:** San Francisco, CA
**Role:** AI Research
**Salary:** $350k – $475k/yr
**Skills:** Python, PyTorch, TensorFlow, JAX, RLHF, Rlaif, Preference Modeling, Safety Evaluations, Red-Teaming, Synthetic Data, Distributed Training, Jailbreaking, Scalable Oversight, Reward Hacking, Agentic Tasks
**Posted:** 2026-08-24

> Conduct AI safety research across data curation, post-training, evaluations, synthetic data, and red-teaming to improve model reliability on harmful and dual-use requests. The role requires AI safety experience, Python, deep learning frameworks, and scalable technical research skills.

## Job Description

## Responsibilities
- Investigate how models handle harmful, sensitive, and dual-use requests, including what they learn from data and how training shapes reliable refusal and engagement boundaries.
- Build data-filtering pipelines and quality classifiers for pre-training corpora, and study downstream safety effects.
- Apply post-training methods, including **RLHF**, **RLAIF**, and policy-based reasoning approaches.
- Design, build, and maintain safety evaluations for long-horizon and agentic tasks.
- Generate and curate synthetic data for training and evaluating refusal boundaries and safety-relevant behaviors.
- Red-team models and products to identify failure modes, jailbreaks, and emergent risks, then design mitigations.

## Requirements
- Bachelor's degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline.
- Background in AI safety research and hands-on experience with at least one of RLHF/RLAIF, alignment and preference modeling, deliberative alignment, safety evaluations, or red-teaming.
- Proficiency in **Python** and familiarity with deep learning frameworks such as **PyTorch**, **TensorFlow**, or **JAX**.
- Experience debugging distributed training and writing scalable code.
- Strong written communication and ability to explain complex technical concepts.

## Nice-to-haves
- Experience evaluating long-horizon, multi-step, or agentic tasks.
- Experience generating synthetic data at scale.
- Experience with modern red-teaming and jailbreaking techniques.
- AI safety research contributions, including publications, open-source evaluations, or public red-teaming work.
- Familiarity with scalable oversight, reward hacking, and jailbreak robustness.
- PhD in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline, or equivalent industry research experience.

## Compensation & Benefits
- Annual salary range: **$350,000–$475,000 USD**.
- Health, dental, and vision benefits.
- Unlimited PTO.
- Paid parental leave.
- Relocation support as needed.
- Visa sponsorship available.

## Similar jobs

- [Research Engineer, Takeoff Intel](https://hotfix.jobs/jobs/398824a2-65cc-4e28-aaeb-26c3b6610876) - Anthropic - San Francisco, CA - $350k – $850k/yr
- [Research, Coding Agents](https://hotfix.jobs/jobs/53bac391-9bcb-48e0-84f9-2dbe7cd426b8) - Thinking Machines Lab - San Francisco, CA - $350k – $475k/yr
- [Researcher, Agent Safety, Oversight and System Mitigations](https://hotfix.jobs/jobs/4544e3bb-bb96-43d2-a96c-cd364b641660) - OpenAI - San Francisco, CA - $380k – $500k/yr
- [Researcher, Agent Safety, Training and Evaluations](https://hotfix.jobs/jobs/d80336da-e453-4999-9f26-85a125b679d9) - OpenAI - San Francisco, CA - $380k – $500k/yr
- [Research Scientist, Life Sciences](https://hotfix.jobs/jobs/c9d5ef0c-39c0-4e12-b2b6-d7f30615fb7e) - Anthropic - San Francisco, CA - $300k – $320k/yr

**Apply:** https://hotfix.jobs/jobs/420d2360-0401-41e6-8665-9a5112542a0b
**Canonical:** https://hotfix.jobs/jobs/420d2360-0401-41e6-8665-9a5112542a0b