# Research Scientist / Engineer

**Company:** [Luma AI](https://hotfix.jobs/companies/lumalabs-ai)
**Location:** Redwood City, CA
**Role:** ML Engineering
**Salary:** $188k – $395k/yr
**Experience:** 5+ years
**Skills:** Reinforcement Learning, ppo, RLHF, rlvr, PyTorch, fsdp, vLLM, sglang, Kubernetes, Ray, nccl, mpi
**Posted:** 2026-07-21

> Build and scale distributed reinforcement learning infrastructure for post-training large multimodal foundation models, including rollout generation, environments, rewards, and evaluation systems for agentic tasks.

## Job Description

## Responsibilities
- Design, build, and scale distributed RL post-training systems for large multimodal models — orchestrating trainer, rollout, environment, and reward workloads across thousands of GPUs.
- Build and optimize high-throughput rollout generation, including efficient integration of inference engines (e.g. vLLM, SGLang) into the training loop, weight synchronization, and asynchronous / off-policy training schemes.
- Design and implement RL environments for agentic and multi-step tasks — sandboxed code execution, tool use, computer use, and multimodal interaction — that are reproducible, hermetic, and scalable to millions of episodes.
- Build reward infrastructure: verifiable / programmatic rewards, reward model serving, LLM-as-judge pipelines, and defenses against reward hacking.
- Develop the evaluation, monitoring, and debugging tooling needed to keep large RL runs stable, diagnose convergence and throughput regressions, and understand model behavior mid-run.
- Advance RL training efficiency and stability: sequence packing for long multi-turn trajectories, KV cache reuse across rollouts, curriculum and task sampling, and resource scheduling across heterogeneous training/inference workloads.
- Collaborate closely with researchers to turn new post-training ideas (RLVR, agentic RL, long-horizon credit assignment, self-improvement loops) into production-quality training runs.

## Requirements
- Hands-on experience post-training LLMs with reinforcement learning (e.g. PPO / GRPO-family methods, RLHF, RLVR / RL from verifiable rewards) at meaningful scale.
- Extensive experience with distributed PyTorch training and parallelization strategies (FSDP, Tensor / Pipeline / Expert Parallel) for foundation models.
- Experience building RL environments, reward functions, verifiers, or evaluation harnesses for LLM agents — including sandboxed execution and multi-turn tool use.
- Deep familiarity with RL post-training frameworks and their systems tradeoffs (e.g. veRL, OpenRLHF, TRL, Ray-based orchestration) and inference engines used for rollouts (vLLM, SGLang).
- Strong understanding of GPU clusters, networking, and communication libraries (NCCL, MPI), and how they behave under mixed training + inference workloads.

## Nice-to-Haves
- Experience running RL training across >100 GPUs, including asynchronous or disaggregated trainer/rollout architectures.
- Experience with containerization and orchestration (Kubernetes, Ray) for large environment fleets and sandboxed workloads.
- Research contributions in RL for LLMs — reasoning, agents, reward modeling, or long-horizon tasks — or open-source contributions to RL training frameworks.

## Compensation
- Base pay range: $187,500 – $395,000 per year.

## Similar roles

- [Software Engineer, Inference](https://hotfix.jobs/jobs/c31a9085-9d3a-4515-b43b-42516d45d8f1) - Luma AI - Palo Alto, CA - $188k – $395k/yr
- [Software Engineer, ML Platform](https://hotfix.jobs/jobs/a9f2a8c2-33a8-4d0d-91c4-60f6a5135e7d) - Luma AI - Palo Alto, CA - $188k – $395k/yr
- [Research Scientist / Engineer – Training Infrastructure](https://hotfix.jobs/jobs/df564675-5da1-475c-a6b4-e81cc62da4eb) - Luma AI - Palo Alto, CA - $188k – $395k/yr
- [Software Engineer, Machine Learning Platform](https://hotfix.jobs/jobs/8227a48c-d3ba-4a85-afb5-a934accf63b3) - Chime - San Francisco, CA - $187k – $259k/yr
- [Agentic AI Engineer](https://hotfix.jobs/jobs/d74d66cb-3089-4707-b593-9f9e90c11b05) - EigenLayer - Remote - $187k – $253k/yr

**Apply:** https://hotfix.jobs/jobs/063941d9-e414-4af2-b6c8-dff6c2af103c
**Canonical:** https://hotfix.jobs/jobs/063941d9-e414-4af2-b6c8-dff6c2af103c