ML Researcher - Posttraining
Conduct research and engineering on large-scale post-training of diffusion and language models, focusing on aesthetics, preference optimization, reinforcement learning, reward modeling, and evaluation. The role requires strong PyTorch, distributed training, low-precision inference, and VLM experience.
About the job
Responsibilities
- Fine-tune diffusion models at scale to improve image aesthetics and quality.
- Implement post-training techniques including supervised fine-tuning, preference optimization, reinforcement learning, on-policy distillation, and distillation/acceleration methods.
- Design evaluation suites and reward functions for image-space reinforcement learning.
- Train custom vision-language models (VLMs) as reward models.
- Fine-tune custom large language models (LLMs) for prompt expansion using reinforcement learning.
- Coordinate preference-data collection and model-evaluation results with data teams and partners.
- Work on safety alignment for open-source model releases.
- Collaborate with AI research and engineering teams to integrate research advances into products.
Requirements
- Demonstrated experience post-training diffusion models for image or video generation.
- Experience with large-scale model training, inference, and optimization.
- Strong understanding of LLM and diffusion post-training pipelines and algorithms, including PPO, GRPO, DPO, OPD, and MOPD.
- Strong proficiency in PyTorch and understanding of its internals.
- Experience with distributed training paradigms including FSDP, context parallelism, sequence parallelism, USP, tensor parallelism, and expert parallelism.
- Knowledge of low-precision training and inference, including FP8, NVFP4, and MXFP8.
- Understanding of fast inference engines such as vLLM and sglang, and LLM reinforcement-learning frameworks such as slime, miles, tinker, and verl.
- Understanding of reinforcement-learning infrastructure and optimization techniques including asynchronous RL, fast weight transfer, rollout pipelining, and off-policy data management.
- Experience training VLMs.
- Ability to monitor model regressions, identify weak areas, and translate them into evaluations and reward designs.
- Comfortable working in a goal-oriented, ambiguous research environment and turning open-ended goals into concrete plans and execution items.
- Good judgment in selecting and scaling training strategies across compute and data.
Nice to Have
- Ongoing engagement with research in LLMs, VLMs, representation learning, or robotics.
- Strong research taste, with a preference for simple methods that scale with compute and data while minimizing human supervision.
Compensation and Benefits
- Competitive compensation with salary and equity packages.
- Health, dental, and vision insurance premiums covered for employees.
- Health FSA and long-term disability coverage.
- Flexible paid time off.
- 401(k) with a 4% company-sponsored match.
- Covered office meals and transit to and from the office.
- Potential international visa sponsorship.
Skills
Diffusion Models, PyTorch, Ppo, Grpo, Dpo, Reinforcement Learning, Distributed Training, Fsdp, Tensor Parallelism, Fp8, vLLM, Vlms, LLMs, Reward Modeling, Model Fine-Tuning
Similar jobs
AI Research jobsConduct applied research on AI agents, designing experiments and evaluation systems to improve reliability, context retention, and multi-step task completion. The role requires strong AI/ML research, engineering, experimental design, and communication skills.
Research Engineer building large-scale AI capability evaluations, telemetry, data pipelines, and analysis tools for Anthropic’s Takeoff Intel team. The role requires hands-on large language model experimentation, rapid prototyping, data expertise, and strong research collaboration.
Conduct applied research on foundation models for fraud detection using large-scale behavioral and financial-risk data. The role spans experimentation, evaluation, production deployment, and cross-functional work on model governance, requiring 4+ years of applied ML experience and strong Python and SQL skills.
Researcher or engineer focused on designing, evaluating, and productionizing oversight systems and safety mitigations for autonomous AI agents. The role requires strong systems or security reasoning, threat-modeling ability, and experience building practical evaluations and controls.
Researcher focused on training and evaluating frontier AI agents, mining incidents, and building scalable safety measurement systems. The role requires strong research or ML engineering execution, quantitative judgment, and the ability to own ambiguous projects end to end.