Skip to content

Research, RL Scaling

Researcher focused on scaling reinforcement learning for frontier models, with ownership spanning asynchronous RL algorithms, inference and distributed training systems, and large-scale empirical studies. Requires strong Python and deep learning experience, scalable systems debugging, and rigorous research judgment.

About the job

Responsibilities

  • Co-design reinforcement learning recipes and the systems that run them at frontier scale.
  • Advance asynchronous reinforcement learning algorithms.
  • Improve rollout-generation efficiency and integrate inference with training.
  • Run frontier-scale reinforcement learning end to end, including bringing up models and training setups and maintaining run stability.
  • Optimize accelerator utilization, memory, communication, and low-precision numerics.
  • Conduct ablations and scaling studies with reliable instrumentation and clear write-ups.

Requirements

  • Proficiency in Python and familiarity with at least one deep learning framework, such as PyTorch, TensorFlow, or JAX.
  • Experience debugging distributed training and writing scalable code.
  • Bachelor’s degree or equivalent experience in computer science, machine learning, physics, mathematics, or a related discipline.
  • Strong written communication and ability to explain complex technical concepts.
  • Strong research judgment, including clean ablations, honest baselines, and clear technical writing.

Nice-to-haves

  • PhD or equivalent industry research experience.
  • Strong grounding in reinforcement learning for large language models and modern policy optimization methods.
  • Deep understanding of asynchronous reinforcement learning algorithms and ML/systems trade-offs.
  • Experience training large models across many accelerators and working with distributed parallelism, memory, and communication.
  • Knowledge of inference systems, rollout throughput, and cost optimization.
  • Experience building or operating decoupled generation/training reinforcement learning systems at scale.
  • Experience with verifiable and agentic tasks, including multi-turn environments.
  • Experience with reinforcement learning training stability techniques.
  • Familiarity with low-precision training and inference and quantization.
  • Hands-on experience with LLM serving stacks such as SGLang, vLLM, TokenSpeed, or custom engines.
  • Experience with large-model scaling studies.
  • Contributions to open-source training or inference frameworks.

Compensation and Benefits

  • Annual salary range: $350,000–$475,000 USD.
  • Visa sponsorship.
  • Health, dental, and vision benefits.
  • Unlimited paid time off.
  • Paid parental leave.
  • Relocation support as needed.

Skills

Python, PyTorch, TensorFlow, JAX, Reinforcement Learning, Asynchronous Reinforcement Learning, LLMs, Distributed Training, Parallelism Strategies, Inference Systems, Quantization, vLLM, Sglang, Low-Precision Numerics

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Research Software Engineer, Post Training
$350k+/yrHybridML Engineering

Build and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

AI Infrastructure Engineer
$350k+/yrOn-site4+ YOEML Engineering

Operates and improves the infrastructure powering large-scale post-training and reinforcement learning runs, partnering with researchers to debug failures, improve reliability, and automate recovery. Requires 4+ years operating distributed production systems and strong Python, Go, or C++ skills.

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Research, General Agents
$350k+/yrHybridML Engineering

Research-focused engineer advancing agentic model capabilities across synthetic data, task environments, evaluations, training, and usability improvements. Requires strong Python engineering, deep learning framework experience, scalable distributed training skills, and scientific experimentation ability.

OpenAI

OpenAI

San Francisco, CA

Machine Learning Engineer, Multimodal Perception and Authentication
$342k+/yrHybridML Engineering

Develop multimodal perception and authentication systems combining visual, audio, and other sensor signals for real-world AI products. The role requires machine learning expertise, practical research-to-system experience, and proficiency in Python and PyTorch with comfort in C++.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Trainium
$295k+/yrHybrid3+ YOEML Engineering

Build and optimize OpenAI’s inference stack for AWS Trainium across high-performance kernels, compilers, runtimes, and model execution. The role requires systems programming and accelerator experience, with opportunities to solve end-to-end performance problems for frontier-scale AI models.