Skip to content

Member of Technical Staff - Research Engineer

Research Engineer focused on optimizing and stabilizing large-scale GPU training for multimodal generative models. The role combines low-level kernel and precision work, distributed training debugging, profiling, benchmarking, and close collaboration with researchers.

About the job

Responsibilities

  • Improve the performance, reliability, and numerical stability of production training runs for large multimodal generative models.
  • Profile full training steps across model code, attention, kernels, data loading, encoders, communication, optimizer steps, checkpointing, and memory pressure.
  • Implement and validate GPU-level optimizations, including fused kernels, attention paths, low-precision matrix multiplications, quantization kernels, and CUDA/Triton/CuTe/CUTLASS experiments.
  • Advance lower-precision training using FP8, MXFP8, FP4-style paths, weight and activation quantization, accumulation strategies, numerical validation, and convergence analysis.
  • Translate architecture changes into efficient training implementations with researchers and distinguish model-quality improvements from microbenchmark gains.
  • Debug distributed training failures such as NaNs, loss spikes, numerical drift, memory leaks, stragglers, faulty nodes, NCCL issues, and throughput cliffs.
  • Build benchmarking and profiling harnesses across hardware, shapes, sequence lengths, and training configurations.
  • Turn recurring failures into improved abstractions and tools.

Requirements

  • Deep experience with large-scale training systems, preferably working closely with research teams.
  • Strong PyTorch fluency and ability to modify low-level training code.
  • Experience with distributed training concepts including FSDP, tensor/model/context/sequence parallelism, activation checkpointing, NCCL, and overlapping computation and communication.
  • Hands-on experience improving throughput, memory footprint, or stability in real training runs.
  • Experience profiling GPU workloads with tools such as Nsight Systems, Nsight Compute, torch profiler, trace viewers, or custom telemetry.
  • Ability to verify correctness, numerical behavior, and performance of GPU and training-system changes.
  • Understanding of low-precision training and quantization tradeoffs, including scaling, accumulation, numerical validation, and convergence risk.
  • Strong research judgment and ability to connect optimization work to model-quality outcomes.
  • Comfort working independently on ambiguous implementation, debugging, and investigation tasks.

Nice-to-haves

  • Supported or co-owned training for a frontier foundation model that shipped or reached a major release.
  • Written or substantially improved forward/backward GPU kernels.
  • Experience with attention performance, variable sequence length training, or non-standard attention patterns.
  • Experience with Hopper- or Blackwell-class GPUs.
  • Experience with low-precision training.
  • Experience with diffusion, flow matching, DiT, and multimodal generative model training.
  • Ability to move between profiler traces, kernel code, distributed-systems failures, and research discussions.

Compensation

  • Europe: €130,000–€240,000 base annual salary plus equity.
  • United States: $180,000–$290,000 base annual salary plus equity.

Skills

PyTorch, Fsdp, CUDA, Triton, Cute, Cutlass, Nccl, Fp8, Mxfp8, Quantization, Nsight Systems, Nsight Compute, Diffusion, Flow Matching, Dit

Black Forest Labs

Black Forest Labs

Freiburg, Germany

Member of Technical Staff - Image / Video Generation
€130k+/yrHybridML Engineering

Trains and fine-tunes large-scale diffusion transformer models for image and video generation, conducts rigorous ablation studies, and optimizes distributed training. Requires hands-on diffusion-model experience, strong PyTorch and transformer expertise, and understanding of generative-model evaluation.

Mintlify

Mintlify

San Francisco, CA

Applied AI Engineer
$130k+/yrOn-site4+ YOEML Engineering

Build and own customer-facing AI products from experimentation through production, including reliable agents, evaluation systems, APIs, interfaces, and infrastructure. Requires at least four years of software development experience and deep production experience with language-model systems.

PathAI

PathAI

Boston, MA
Machine Learning Engineer III
$131k+/yrOn-site5+ YOEML Engineering

Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.

Build

Build

New York, NY
AI Engineer (Assistant)
$125k+/yrOn-siteML Engineering

Build and ship production agentic AI workflows for complex real estate and built-world processes. The role combines product engineering, applied AI, customer collaboration, workflow orchestration, evaluation, and reliable user-facing experiences.

Fab2

Fab2

Austin, TX
Software Engineer, AI Platform
$140k+/yrOn-siteML Engineering

Build the AI platform behind fab2, including model infrastructure, agent systems, evaluations, and tools for engineering and fab operations. The role requires strong production software engineering skills and comfort working across frontend, backend, infrastructure, and data.