# Member of Technical Staff - Research Engineer

**Company:** [Black Forest Labs](https://hotfix.jobs/companies/black-forest-labs)
**Location:** San Francisco, CA, Freiburg, Germany
**Role:** ML Engineering
**Salary:** €130k – €240k/yr
**Skills:** PyTorch, Fsdp, CUDA, Triton, Cute, Cutlass, Nccl, Fp8, Mxfp8, Quantization, Nsight Systems, Nsight Compute, Diffusion, Flow Matching, Dit
**Posted:** 2026-06-29

> Research Engineer focused on optimizing and stabilizing large-scale GPU training for multimodal generative models. The role combines low-level kernel and precision work, distributed training debugging, profiling, benchmarking, and close collaboration with researchers.

## Job Description

## Responsibilities
- Improve the performance, reliability, and numerical stability of production training runs for large multimodal generative models.
- Profile full training steps across model code, attention, kernels, data loading, encoders, communication, optimizer steps, checkpointing, and memory pressure.
- Implement and validate GPU-level optimizations, including fused kernels, attention paths, low-precision matrix multiplications, quantization kernels, and CUDA/Triton/CuTe/CUTLASS experiments.
- Advance lower-precision training using FP8, MXFP8, FP4-style paths, weight and activation quantization, accumulation strategies, numerical validation, and convergence analysis.
- Translate architecture changes into efficient training implementations with researchers and distinguish model-quality improvements from microbenchmark gains.
- Debug distributed training failures such as NaNs, loss spikes, numerical drift, memory leaks, stragglers, faulty nodes, NCCL issues, and throughput cliffs.
- Build benchmarking and profiling harnesses across hardware, shapes, sequence lengths, and training configurations.
- Turn recurring failures into improved abstractions and tools.

## Requirements
- Deep experience with large-scale training systems, preferably working closely with research teams.
- Strong PyTorch fluency and ability to modify low-level training code.
- Experience with distributed training concepts including FSDP, tensor/model/context/sequence parallelism, activation checkpointing, NCCL, and overlapping computation and communication.
- Hands-on experience improving throughput, memory footprint, or stability in real training runs.
- Experience profiling GPU workloads with tools such as Nsight Systems, Nsight Compute, torch profiler, trace viewers, or custom telemetry.
- Ability to verify correctness, numerical behavior, and performance of GPU and training-system changes.
- Understanding of low-precision training and quantization tradeoffs, including scaling, accumulation, numerical validation, and convergence risk.
- Strong research judgment and ability to connect optimization work to model-quality outcomes.
- Comfort working independently on ambiguous implementation, debugging, and investigation tasks.

## Nice-to-haves
- Supported or co-owned training for a frontier foundation model that shipped or reached a major release.
- Written or substantially improved forward/backward GPU kernels.
- Experience with attention performance, variable sequence length training, or non-standard attention patterns.
- Experience with Hopper- or Blackwell-class GPUs.
- Experience with low-precision training.
- Experience with diffusion, flow matching, DiT, and multimodal generative model training.
- Ability to move between profiler traces, kernel code, distributed-systems failures, and research discussions.

## Compensation
- Europe: €130,000–€240,000 base annual salary plus equity.
- United States: $180,000–$290,000 base annual salary plus equity.

## Similar jobs

- [Member of Technical Staff - Image / Video Generation](https://hotfix.jobs/jobs/405a95a1-d8ff-49a2-9851-5a5e1cf05ad8) - Black Forest Labs - Freiburg, Germany - €130k – €340k/yr
- [Applied AI Engineer](https://hotfix.jobs/jobs/2f2e7873-cc6a-476d-a418-7b067cadf604) - Mintlify - San Francisco, CA - $130k – $200k/yr
- [Machine Learning Engineer III](https://hotfix.jobs/jobs/2359be26-5005-4fa7-94c9-8a86066a6bb5) - PathAI - Boston, MA - $131k – $200k/yr
- [AI Engineer (Assistant)](https://hotfix.jobs/jobs/4c7ec3aa-a4bf-4dc9-987f-c541087f251b) - Build - New York, NY - $125k – $225k/yr
- [Software Engineer, AI Platform](https://hotfix.jobs/jobs/7dcee5ac-38bb-4da0-b399-0b5896976722) - Fab2 - Austin, TX - $140k – $200k/yr

**Apply:** https://hotfix.jobs/jobs/ff508d2d-9bb6-4284-870a-db963445dd59
**Canonical:** https://hotfix.jobs/jobs/ff508d2d-9bb6-4284-870a-db963445dd59