# ML Researcher - Image / Video Diffusion

**Company:** [Krea](https://hotfix.jobs/companies/krea)
**Location:** San Francisco, CA
**Role:** AI Research
**Skills:** PyTorch, Diffusion Models, Distributed Training, Fsdp, Tensor Parallelism, Pipeline Parallelism, Expert Parallelism, Gpu Clusters, Nccl, InfiniBand, Nvlink, Fp8, Nvfp4, Mxfp8, Reinforcement Learning
**Posted:** 2026-09-01

> Researcher developing and scaling image and video diffusion models on large GPU clusters. The role requires deep PyTorch and distributed-training expertise, proficiency in low-precision computation, and experience profiling and debugging large-scale model training.

## Job Description

## Responsibilities
- Train diffusion models for image and video generation on large GPU clusters.
- Optimize and profile large distributed training runs across model architectures, kernels, data loading, memory constraints, and communication.
- Implement and improve distributed training strategies including FSDP, CP, SP, TP, and EP.
- Improve model quality and reliability through data, model architecture, training pipelines, experiment structure, and evaluation design.
- Debug distributed training errors and implement fault-tolerance solutions, including identifying faulty GPU, NVLink, and InfiniBand components and monitoring numerical and NCCL issues.
- Run ablations across architecture, attention, optimizer, data, and algorithmic choices to improve model efficiency and performance.
- Design custom data pipelines to improve data quality and translate ambiguous research goals into concrete plans and experiments.

## Requirements
- Proven experience working with image or video models at scale; publications or open-source contributions are a plus.
- Strong proficiency in PyTorch and understanding of its internals.
- Strong background in distributed training paradigms including FSDP, CP, SP, USP, TP, and EP, including their tradeoffs.
- Experience profiling and debugging large distributed training runs and analyzing traces to identify bottlenecks.
- Knowledge of low-precision training and inference with FP8, NVFP4, and MXFP8.
- Solid understanding of diffusion model training across pretraining, midtraining, preference optimization, and reinforcement learning.
- Ability to work independently in a goal-oriented research environment, handle underspecified goals, iterate rapidly, and propose creative research directions.
- Good judgment about selecting and scaling training strategies, compute, and data.
- Strong research taste, with a bias toward simple, scalable methods requiring minimal human supervision.

## Nice-to-haves
- Publications or open-source contributions related to image or video models.
- Familiarity with developments in LLM, VLM, representation learning, and robotics research.

## Compensation and Benefits
- Competitive compensation with generous salary and equity packages.
- 100% employee health premium coverage and 99% dental and vision premium coverage.
- Health FSA and long-term disability coverage.
- Flexible PTO.
- 401(k) with a 4% company-sponsored match.
- Covered office meals and Uber transportation to and from the office.
- Visa sponsorship may be available for eligible international candidates.

## Similar jobs

- [AI Researcher](https://hotfix.jobs/jobs/077574a5-3c6d-4634-a23c-4a909ce8aa65) - Improbable - Remote
- [Research Engineer, Takeoff Intel](https://hotfix.jobs/jobs/398824a2-65cc-4e28-aaeb-26c3b6610876) - Anthropic - San Francisco, CA - $350k – $850k/yr
- [Applied AI Research Scientist](https://hotfix.jobs/jobs/93baef6f-91a8-4c62-acaa-44c3ea48b467) - Sardine - Remote
- [Researcher, Agent Safety, Oversight and System Mitigations](https://hotfix.jobs/jobs/4544e3bb-bb96-43d2-a96c-cd364b641660) - OpenAI - San Francisco, CA - $380k – $500k/yr
- [Researcher, Agent Safety, Training and Evaluations](https://hotfix.jobs/jobs/d80336da-e453-4999-9f26-85a125b679d9) - OpenAI - San Francisco, CA - $380k – $500k/yr

**Apply:** https://hotfix.jobs/jobs/2a25c729-b3b2-4817-be79-67e801b55f23
**Canonical:** https://hotfix.jobs/jobs/2a25c729-b3b2-4817-be79-67e801b55f23