# ML Researcher - Posttraining

**Company:** [Krea](https://hotfix.jobs/companies/krea)
**Location:** San Francisco, CA
**Role:** AI Research
**Skills:** Diffusion Models, PyTorch, Ppo, Grpo, Dpo, Reinforcement Learning, Distributed Training, Fsdp, Tensor Parallelism, Fp8, vLLM, Vlms, LLMs, Reward Modeling, Model Fine-Tuning
**Posted:** 2026-09-01

> Conduct research and engineering on large-scale post-training of diffusion and language models, focusing on aesthetics, preference optimization, reinforcement learning, reward modeling, and evaluation. The role requires strong PyTorch, distributed training, low-precision inference, and VLM experience.

## Job Description

## Responsibilities
- Fine-tune diffusion models at scale to improve image aesthetics and quality.
- Implement post-training techniques including supervised fine-tuning, preference optimization, reinforcement learning, on-policy distillation, and distillation/acceleration methods.
- Design evaluation suites and reward functions for image-space reinforcement learning.
- Train custom vision-language models (VLMs) as reward models.
- Fine-tune custom large language models (LLMs) for prompt expansion using reinforcement learning.
- Coordinate preference-data collection and model-evaluation results with data teams and partners.
- Work on safety alignment for open-source model releases.
- Collaborate with AI research and engineering teams to integrate research advances into products.

## Requirements
- Demonstrated experience post-training diffusion models for image or video generation.
- Experience with large-scale model training, inference, and optimization.
- Strong understanding of LLM and diffusion post-training pipelines and algorithms, including PPO, GRPO, DPO, OPD, and MOPD.
- Strong proficiency in PyTorch and understanding of its internals.
- Experience with distributed training paradigms including FSDP, context parallelism, sequence parallelism, USP, tensor parallelism, and expert parallelism.
- Knowledge of low-precision training and inference, including FP8, NVFP4, and MXFP8.
- Understanding of fast inference engines such as vLLM and sglang, and LLM reinforcement-learning frameworks such as slime, miles, tinker, and verl.
- Understanding of reinforcement-learning infrastructure and optimization techniques including asynchronous RL, fast weight transfer, rollout pipelining, and off-policy data management.
- Experience training VLMs.
- Ability to monitor model regressions, identify weak areas, and translate them into evaluations and reward designs.
- Comfortable working in a goal-oriented, ambiguous research environment and turning open-ended goals into concrete plans and execution items.
- Good judgment in selecting and scaling training strategies across compute and data.

## Nice to Have
- Ongoing engagement with research in LLMs, VLMs, representation learning, or robotics.
- Strong research taste, with a preference for simple methods that scale with compute and data while minimizing human supervision.

## Compensation and Benefits
- Competitive compensation with salary and equity packages.
- Health, dental, and vision insurance premiums covered for employees.
- Health FSA and long-term disability coverage.
- Flexible paid time off.
- 401(k) with a 4% company-sponsored match.
- Covered office meals and transit to and from the office.
- Potential international visa sponsorship.

## Similar jobs

- [AI Researcher](https://hotfix.jobs/jobs/077574a5-3c6d-4634-a23c-4a909ce8aa65) - Improbable - Remote
- [Research Engineer, Takeoff Intel](https://hotfix.jobs/jobs/398824a2-65cc-4e28-aaeb-26c3b6610876) - Anthropic - San Francisco, CA - $350k – $850k/yr
- [Applied AI Research Scientist](https://hotfix.jobs/jobs/93baef6f-91a8-4c62-acaa-44c3ea48b467) - Sardine - Remote
- [Researcher, Agent Safety, Oversight and System Mitigations](https://hotfix.jobs/jobs/4544e3bb-bb96-43d2-a96c-cd364b641660) - OpenAI - San Francisco, CA - $380k – $500k/yr
- [Researcher, Agent Safety, Training and Evaluations](https://hotfix.jobs/jobs/d80336da-e453-4999-9f26-85a125b679d9) - OpenAI - San Francisco, CA - $380k – $500k/yr

**Apply:** https://hotfix.jobs/jobs/4548e069-4358-4d58-9598-f58c621bacf0
**Canonical:** https://hotfix.jobs/jobs/4548e069-4358-4d58-9598-f58c621bacf0