# Research, RL Scaling

**Company:** [Thinking Machines Lab](https://hotfix.jobs/companies/thinking-machines-lab)
**Location:** San Francisco, CA
**Role:** ML Engineering
**Salary:** $350k – $475k/yr
**Skills:** Python, PyTorch, TensorFlow, JAX, Reinforcement Learning, Asynchronous Reinforcement Learning, LLMs, Distributed Training, Parallelism Strategies, Inference Systems, Quantization, vLLM, Sglang, Low-Precision Numerics
**Posted:** 2026-08-21

> Researcher focused on scaling reinforcement learning for frontier models, with ownership spanning asynchronous RL algorithms, inference and distributed training systems, and large-scale empirical studies. Requires strong Python and deep learning experience, scalable systems debugging, and rigorous research judgment.

## Job Description

## Responsibilities
- Co-design reinforcement learning recipes and the systems that run them at frontier scale.
- Advance asynchronous reinforcement learning algorithms.
- Improve rollout-generation efficiency and integrate inference with training.
- Run frontier-scale reinforcement learning end to end, including bringing up models and training setups and maintaining run stability.
- Optimize accelerator utilization, memory, communication, and low-precision numerics.
- Conduct ablations and scaling studies with reliable instrumentation and clear write-ups.

## Requirements
- Proficiency in Python and familiarity with at least one deep learning framework, such as PyTorch, TensorFlow, or JAX.
- Experience debugging distributed training and writing scalable code.
- Bachelor’s degree or equivalent experience in computer science, machine learning, physics, mathematics, or a related discipline.
- Strong written communication and ability to explain complex technical concepts.
- Strong research judgment, including clean ablations, honest baselines, and clear technical writing.

## Nice-to-haves
- PhD or equivalent industry research experience.
- Strong grounding in reinforcement learning for large language models and modern policy optimization methods.
- Deep understanding of asynchronous reinforcement learning algorithms and ML/systems trade-offs.
- Experience training large models across many accelerators and working with distributed parallelism, memory, and communication.
- Knowledge of inference systems, rollout throughput, and cost optimization.
- Experience building or operating decoupled generation/training reinforcement learning systems at scale.
- Experience with verifiable and agentic tasks, including multi-turn environments.
- Experience with reinforcement learning training stability techniques.
- Familiarity with low-precision training and inference and quantization.
- Hands-on experience with LLM serving stacks such as SGLang, vLLM, TokenSpeed, or custom engines.
- Experience with large-model scaling studies.
- Contributions to open-source training or inference frameworks.

## Compensation and Benefits
- Annual salary range: $350,000–$475,000 USD.
- Visa sponsorship.
- Health, dental, and vision benefits.
- Unlimited paid time off.
- Paid parental leave.
- Relocation support as needed.

## Similar jobs

- [Research Software Engineer, Post Training](https://hotfix.jobs/jobs/168e3c8f-8577-4482-bf94-91b3d11744ba) - Thinking Machines Lab - San Francisco, CA - $350k – $475k/yr
- [AI Infrastructure Engineer](https://hotfix.jobs/jobs/0d5aa4cf-a861-427b-8d8f-cba8ac94104e) - Thinking Machines Lab - San Francisco, CA - $350k – $475k/yr
- [Research, General Agents](https://hotfix.jobs/jobs/e075e229-9ae8-46ba-92a5-20a0a6d4f0db) - Thinking Machines Lab - San Francisco, CA - $350k – $475k/yr
- [Machine Learning Engineer, Multimodal Perception and Authentication](https://hotfix.jobs/jobs/26314e6e-423f-4a3f-b400-ee0b56429f64) - OpenAI - San Francisco, CA - $342k – $399k/yr
- [Software Engineer, Trainium](https://hotfix.jobs/jobs/5e1a8341-f0e9-45c8-8492-12782b38f079) - OpenAI - San Francisco, CA - $295k – $380k/yr

**Apply:** https://hotfix.jobs/jobs/bd1b548c-7c82-4627-84a9-149150dad06d
**Canonical:** https://hotfix.jobs/jobs/bd1b548c-7c82-4627-84a9-149150dad06d