Skip to content

Member of Technical Staff — RL Research

New/recent PhD to own RL and post-training for large-scale omni models. Build and scale the full RL/post-training stack including rollout, optimization, reward modeling, and evaluation for real-time audiovisual AI.

About the job

What You’ll Own

  • Build Nuance’s RL/post-training stack from 0→1: rollout generation, policy optimization, reward/reference model serving, data feedback loops, evaluation, checkpointing, observability, and debugging.
  • Develop and scale post-training methods such as PPO, GRPO, DPO, rejection sampling, RLHF/RLAIF, online RL, and model-based data improvement.
  • Design the systems abstractions that connect research ideas to production-scale RL runs: trainers, rollout workers, reward models, evaluators, data queues, experience buffers, and checkpoint promotion.
  • Build evaluation and feedback loops for omni behavior: turn-taking, interruption, timing, emotional response, audiovisual coherence, instruction following, and real-time interaction quality.
  • Optimize the end-to-end post-training loop across rollout throughput, serving latency, GPU utilization, policy update efficiency, queueing, checkpoint overhead, and research iteration speed.
  • Evolve the platform as algorithms, model architectures, reward definitions, data sources, and evaluation methods change.

What We’re Looking For

  • A PhD — completed, or in its final stretch — in ML, RL, or a related field, with research depth shown through publications, a strong lab/advisor, or substantial open-source work.
  • Solid understanding of RL/post-training methods: policy optimization, reward modeling, preference optimization, rejection sampling, KL control, evaluation, and data feedback loops.
  • Ability to reason about model behavior and training dynamics: reward hacking, unstable rewards, distribution shift, stale policies, mode collapse, over-optimization, noisy preferences, and evaluation mismatch.
  • Exposure to RL/post-training pipelines through research, internships, or open-source — with frameworks such as verl, ms-swift, OpenRLHF, or equivalent, and familiarity with rollout serving systems such as vLLM.
  • Strong software engineering fundamentals and the appetite to build real systems, not just prototypes.
  • Curiosity and adaptability toward new RL algorithms, model architectures, serving systems, evaluation methods, and research ideas.

Bonus Points

  • Hands-on experience with omni or multimodal post-training for audio-video-language models, especially long-context or real-time interactive systems.
  • Experience with PPO, GRPO, DPO, online RL, RLHF/RLAIF, reward modeling, preference data, synthetic data generation, or model-based data improvement.
  • Prior 0→1 experience building post-training systems, RL pipelines, agent training systems, evaluation platforms, or model improvement loops.
  • Experience with adjacent areas such as distributed pretraining, data infrastructure, inference serving, simulation, human/AI feedback collection, or evaluation infrastructure.
  • Publications or substantial open-source contributions in RL, post-training, alignment, evaluation, ML systems, or model behavior.

Compensation

  • $250,000 – $350,000 base salary, plus meaningful equity.

Benefits

  • HSA plan with ~$2,000 in annual company contributions.
  • 15 days of PTO plus public holidays, and office closure for a full week at year-end.
  • Lunch, drinks, and snacks provided every workday.
  • Commuter benefits.
  • 401(k) in progress.

Skills

Reinforcement Learning, Ppo, Dpo, RLHF, Reward Modeling, vLLM, Verl, Openrlhf, Policy Optimization, Distributed Training

Stripe

Stripe

South San Francisco, CA

Machine Learning Engineer
$212k+/yrHybrid2+ YOEML Engineering

Build and productionize scalable machine learning models and systems for underwriting and portfolio management. The role requires a bachelor's degree and at least two years of experience shipping ML systems, plus expertise in model development, deployment, data pipelines, and deep learning.

AfterQuery

AfterQuery

San Francisco, CA

Research Engineer
$210k+/yrOn-site2+ YOEML Engineering

Research Engineer designing post-training infrastructure and running controlled experiments to measure how datasets affect foundation-model behavior. Requires at least 2 years of ML or research engineering experience, strong Python, and hands-on experience with PyTorch, JAX, Ray, Slurm, and LLM post-training.

Decagon

Decagon

San Francisco, CA
Research Engineer, Audio and Speech
$200k+/yrOn-site2+ YOEML Engineering

Research Engineer focused on building and deploying real-time audio and speech models for conversational voice agents. The role requires experience with speech or multimodal machine learning, production inference, Python, and PyTorch, with emphasis on taking research from prototype to measurable production impact.

Decagon

Decagon

San Francisco, CA
Research Engineer, Safety
$200k+/yrOn-site2+ YOEML Engineering

Research Engineer focused on making conversational AI agents safe, reliable, and controllable in production. The role develops evaluations, safeguards, post-training methods, and monitoring systems, requiring 2+ years of AI/ML or safety experience and strong Python and production engineering skills.

Earnin

Earnin

Mountain View, CA

Machine Learning Engineer
$187k+/yrHybrid2+ YOEML Engineering

Machine learning engineer who trains, evaluates, and productionizes models and LLM-powered applications for financial products. Requires 2+ years of ML systems experience, strong Python and PyTorch skills, production data pipelines, model evaluation, and API development.