Skip to content
KreaKrea

ML Researcher - Image / Video Diffusion

Researcher developing and scaling image and video diffusion models on large GPU clusters. The role requires deep PyTorch and distributed-training expertise, proficiency in low-precision computation, and experience profiling and debugging large-scale model training.

About the job

Responsibilities

  • Train diffusion models for image and video generation on large GPU clusters.
  • Optimize and profile large distributed training runs across model architectures, kernels, data loading, memory constraints, and communication.
  • Implement and improve distributed training strategies including FSDP, CP, SP, TP, and EP.
  • Improve model quality and reliability through data, model architecture, training pipelines, experiment structure, and evaluation design.
  • Debug distributed training errors and implement fault-tolerance solutions, including identifying faulty GPU, NVLink, and InfiniBand components and monitoring numerical and NCCL issues.
  • Run ablations across architecture, attention, optimizer, data, and algorithmic choices to improve model efficiency and performance.
  • Design custom data pipelines to improve data quality and translate ambiguous research goals into concrete plans and experiments.

Requirements

  • Proven experience working with image or video models at scale; publications or open-source contributions are a plus.
  • Strong proficiency in PyTorch and understanding of its internals.
  • Strong background in distributed training paradigms including FSDP, CP, SP, USP, TP, and EP, including their tradeoffs.
  • Experience profiling and debugging large distributed training runs and analyzing traces to identify bottlenecks.
  • Knowledge of low-precision training and inference with FP8, NVFP4, and MXFP8.
  • Solid understanding of diffusion model training across pretraining, midtraining, preference optimization, and reinforcement learning.
  • Ability to work independently in a goal-oriented research environment, handle underspecified goals, iterate rapidly, and propose creative research directions.
  • Good judgment about selecting and scaling training strategies, compute, and data.
  • Strong research taste, with a bias toward simple, scalable methods requiring minimal human supervision.

Nice-to-haves

  • Publications or open-source contributions related to image or video models.
  • Familiarity with developments in LLM, VLM, representation learning, and robotics research.

Compensation and Benefits

  • Competitive compensation with generous salary and equity packages.
  • 100% employee health premium coverage and 99% dental and vision premium coverage.
  • Health FSA and long-term disability coverage.
  • Flexible PTO.
  • 401(k) with a 4% company-sponsored match.
  • Covered office meals and Uber transportation to and from the office.
  • Visa sponsorship may be available for eligible international candidates.

Skills

PyTorch, Diffusion Models, Distributed Training, Fsdp, Tensor Parallelism, Pipeline Parallelism, Expert Parallelism, Gpu Clusters, Nccl, InfiniBand, Nvlink, Fp8, Nvfp4, Mxfp8, Reinforcement Learning

Improbable

Improbable

Remote

AI Researcher
No salary listedRemoteAI Research

Conduct applied research on AI agents, designing experiments and evaluation systems to improve reliability, context retention, and multi-step task completion. The role requires strong AI/ML research, engineering, experimental design, and communication skills.

Anthropic

Anthropic

San Francisco, CA

Research Engineer, Takeoff Intel
$350k+/yrHybridAI Research

Research Engineer building large-scale AI capability evaluations, telemetry, data pipelines, and analysis tools for Anthropic’s Takeoff Intel team. The role requires hands-on large language model experimentation, rapid prototyping, data expertise, and strong research collaboration.

Sardine

Sardine

United States

Applied AI Research Scientist
No salary listedRemote4+ YOEAI Research

Conduct applied research on foundation models for fraud detection using large-scale behavioral and financial-risk data. The role spans experimentation, evaluation, production deployment, and cross-functional work on model governance, requiring 4+ years of applied ML experience and strong Python and SQL skills.

OpenAI

OpenAI

San Francisco, CA

Researcher, Agent Safety, Oversight and System Mitigations
$380k+/yrHybridAI Research

Researcher or engineer focused on designing, evaluating, and productionizing oversight systems and safety mitigations for autonomous AI agents. The role requires strong systems or security reasoning, threat-modeling ability, and experience building practical evaluations and controls.

OpenAI

OpenAI

San Francisco, CA

Researcher, Agent Safety, Training and Evaluations
$380k+/yrHybridAI Research

Researcher focused on training and evaluating frontier AI agents, mining incidents, and building scalable safety measurement systems. The role requires strong research or ML engineering execution, quantitative judgment, and the ability to own ambiguous projects end to end.