Research Scientist / Engineer
Build and scale distributed reinforcement learning infrastructure for post-training large multimodal foundation models, including rollout generation, environments, rewards, and evaluation systems for agentic tasks.
About the job
Responsibilities
- Design, build, and scale distributed RL post-training systems for large multimodal models — orchestrating trainer, rollout, environment, and reward workloads across thousands of GPUs.
- Build and optimize high-throughput rollout generation, including efficient integration of inference engines (e.g. vLLM, SGLang) into the training loop, weight synchronization, and asynchronous / off-policy training schemes.
- Design and implement RL environments for agentic and multi-step tasks — sandboxed code execution, tool use, computer use, and multimodal interaction — that are reproducible, hermetic, and scalable to millions of episodes.
- Build reward infrastructure: verifiable / programmatic rewards, reward model serving, LLM-as-judge pipelines, and defenses against reward hacking.
- Develop the evaluation, monitoring, and debugging tooling needed to keep large RL runs stable, diagnose convergence and throughput regressions, and understand model behavior mid-run.
- Advance RL training efficiency and stability: sequence packing for long multi-turn trajectories, KV cache reuse across rollouts, curriculum and task sampling, and resource scheduling across heterogeneous training/inference workloads.
- Collaborate closely with researchers to turn new post-training ideas (RLVR, agentic RL, long-horizon credit assignment, self-improvement loops) into production-quality training runs.
Requirements
- Hands-on experience post-training LLMs with reinforcement learning (e.g. PPO / GRPO-family methods, RLHF, RLVR / RL from verifiable rewards) at meaningful scale.
- Extensive experience with distributed PyTorch training and parallelization strategies (FSDP, Tensor / Pipeline / Expert Parallel) for foundation models.
- Experience building RL environments, reward functions, verifiers, or evaluation harnesses for LLM agents — including sandboxed execution and multi-turn tool use.
- Deep familiarity with RL post-training frameworks and their systems tradeoffs (e.g. veRL, OpenRLHF, TRL, Ray-based orchestration) and inference engines used for rollouts (vLLM, SGLang).
- Strong understanding of GPU clusters, networking, and communication libraries (NCCL, MPI), and how they behave under mixed training + inference workloads.
Nice-to-Haves
- Experience running RL training across >100 GPUs, including asynchronous or disaggregated trainer/rollout architectures.
- Experience with containerization and orchestration (Kubernetes, Ray) for large environment fleets and sandboxed workloads.
- Research contributions in RL for LLMs — reasoning, agents, reward modeling, or long-horizon tasks — or open-source contributions to RL training frameworks.
Compensation
- Base pay range: $187,500 – $395,000 per year.
Skills
Reinforcement Learning, Ppo, RLHF, Rlvr, PyTorch, Fsdp, vLLM, Sglang, Kubernetes, Ray, Nccl, Mpi
Similar jobs
ML Engineering jobsBuild production-grade AI agents, evaluation infrastructure, and developer tooling that make AI-assisted engineering faster, safer, and reusable across teams. The role requires software engineering experience, platform or internal developer-product experience, and hands-on expertise with LLM integration and orchestration.
Build trustworthy infrastructure for production LLM agents, closed-loop evaluation, and autonomous research workflows. The role requires strong Python and distributed-systems experience, hands-on LLM post-training and inference knowledge, and experience operating agent systems at scale.
Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Build and deploy agentic systems that power AI-driven creative video workflows. The role requires 5+ years of experience, production ML or agentic pipeline development, context engineering, and expertise in evaluation and agent infrastructure.
Build and advance agentic machine-learning systems for multimodal creative tasks, with a focus on video understanding, reasoning, control, and tool use. The role requires strong production ML or agent-pipeline experience and deep knowledge of modern LLM techniques.