# Research, Coding Agents

**Company:** [Thinking Machines Lab](https://hotfix.jobs/companies/thinking-machines-lab)
**Location:** San Francisco, CA
**Role:** AI Research
**Salary:** $350k – $475k/yr
**Skills:** Python, PyTorch, TensorFlow, JAX, Reinforcement Learning, Distributed Training, Synthetic Data, Data Pipelines, Reward Modeling, Sandbox Environments, Model Evaluation, Large-Scale Training
**Posted:** 2026-08-27

> Research role focused on improving agentic coding capabilities through reinforcement-learning training, synthetic data, coding environments, reward design, and evaluations. Requires strong Python engineering, scalable distributed-training experience, and a bachelor’s degree or equivalent; research experience and a PhD are preferred.

## Job Description

## Responsibilities
- Design and run reinforcement learning training jobs targeting agentic coding capabilities, iterating on training recipes and data.
- Build and improve sandboxed coding environments and reward signals for model training and evaluation.
- Generate and curate high-quality synthetic coding data and build scalable, general-purpose data pipelines.
- Design evaluations that measure real-world coding usefulness and train models against them to improve day-to-day usability.
- Debug and analyze large-scale reinforcement learning runs to identify confounders, reward hacking, and other failure modes.
- Collaborate with infrastructure, evaluation, and post-training teams on shared data, joint training runs, and usability improvements; ship results into model releases.

## Requirements
- Strong engineering skills, including the ability to contribute code and debug complex codebases.
- Proficiency in Python and familiarity with at least one deep learning framework.
- Experience debugging distributed training and writing scalable code.
- Bachelor’s degree or equivalent experience in computer science, machine learning, physics, mathematics, or a related discipline.
- Clear written communication and ability to explain complex technical concepts.

## Nice-to-Haves
- Experience building synthetic data pipelines and systems adopted and maintained by a team.
- Experience identifying model-usability gaps and addressing them through custom evaluations and training data.
- Experience making large-scale agentic reinforcement learning infrastructure reliable despite failures at scale.
- Experience improving the coding capabilities of a frontier model.
- PhD in computer science, machine learning, physics, mathematics, or a related discipline, or equivalent industry research experience.

## Compensation and Benefits
- Annual salary range: $350,000–$475,000 USD.
- Health, dental, and vision benefits.
- Unlimited paid time off.
- Paid parental leave.
- Relocation support.
- Visa sponsorship.

## Similar jobs

- [Research Engineer, Takeoff Intel](https://hotfix.jobs/jobs/398824a2-65cc-4e28-aaeb-26c3b6610876) - Anthropic - San Francisco, CA - $350k – $850k/yr
- [Research, Safety](https://hotfix.jobs/jobs/420d2360-0401-41e6-8665-9a5112542a0b) - Thinking Machines Lab - San Francisco, CA - $350k – $475k/yr
- [Researcher, Agent Safety, Oversight and System Mitigations](https://hotfix.jobs/jobs/4544e3bb-bb96-43d2-a96c-cd364b641660) - OpenAI - San Francisco, CA - $380k – $500k/yr
- [Researcher, Agent Safety, Training and Evaluations](https://hotfix.jobs/jobs/d80336da-e453-4999-9f26-85a125b679d9) - OpenAI - San Francisco, CA - $380k – $500k/yr
- [Research Scientist, Life Sciences](https://hotfix.jobs/jobs/c9d5ef0c-39c0-4e12-b2b6-d7f30615fb7e) - Anthropic - San Francisco, CA - $300k – $320k/yr

**Apply:** https://hotfix.jobs/jobs/53bac391-9bcb-48e0-84f9-2dbe7cd426b8
**Canonical:** https://hotfix.jobs/jobs/53bac391-9bcb-48e0-84f9-2dbe7cd426b8