Skip to content

Member of Technical Staff - VLM

Advances vision-language models and integrates multimodal capabilities with FLUX diffusion and flow pipelines. The role requires demonstrated VLM pretraining or substantial architectural advancement, strong research or production results, and multi-node distributed training experience.

About the job

Responsibilities

  • Lead development and training of state-of-the-art multimodal vision-language models within the FLUX stack, innovating on architectures rather than only applying existing ones.
  • Design fine-tuning strategies for specialized creative use cases, including captioning, editing instructions, and prompt enhancement.
  • Research integrations between VLM/LLM capabilities and diffusion and flow pipelines to improve generation quality and controllability without computational bottlenecks.
  • Evaluate emerging multimodal architectures and translate recent research into practical improvements.

Requirements

  • Pretrained or significantly advanced a VLM, beyond SFT or LoRA fine-tuning, that was deployed in production or released publicly.
  • Strong publication record or an unambiguous production track record demonstrating work on multimodal architectures.
  • Deep understanding of vision-language representation interaction, including tokenization, alignment, grounding, cross-modal attention, and failure modes.
  • Experience with distributed training at multi-node scale.
  • Comfortable working at the research/production boundary, with a focus on shipping and generalization.

Nice-to-haves

  • Experience with diffusion or flow-based generative models, especially combining autoregressive and diffusion paradigms.

Compensation

  • Base annual salary: €130,000–€340,000
  • Equity

Skills

Vision-Language Models, Multimodal Architectures, LLMs, Diffusion Models, Flow-Based Models, Distributed Training, Multi-Node Training, Tokenization, Cross-Modal Attention, Grounding, Model Fine-Tuning, Python

Black Forest Labs

Black Forest Labs

Freiburg, Germany

Member of Technical Staff - Pretraining
€130k+/yrHybrid7+ YOEAI Research

Leads frontier-scale pretraining research for multimodal image, video, and audio foundation models, shaping architectures, objectives, data strategies, and distributed systems. The role requires prior ownership of production-grade foundation-model pretraining, deep Python and PyTorch expertise, and strong experience with visual generative models.

Vanta

Vanta

Remote

Senior Product Builder, Organizational Intelligence
$176k+/yrRemote5+ YOEAI Research

Build Vanta’s organizational intelligence layer by shipping prototypes, internal tools, and AI agent workflows that make cross-source data useful to EPD, GTM, and other teams. The role requires recent hands-on LLM product work, independent problem scoping, and strong judgment around AI quality, reliability, cost, and latency.

AI Digest

AI Digest

Remote

Research Scientist - Member of Technical Staff
$150k+/yrRemoteAI Research

Conduct research on long-horizon, multi-agent AI behavior by designing agent environments, analyzing large-scale data, and running experiments. The role requires strong research judgment, rapid execution, independence, and familiarity with current AI developments.

AI Digest

AI Digest

Remote

Engineer - Member of Technical Staff
$150k+/yrRemoteAI Research

Build, optimize, and evaluate long-running and multi-agent AI systems, along with tools for monitoring and analyzing their real-world behavior. The role requires software engineering experience with coding agents, strong independence, and familiarity with current AI developments.

Improbable

Improbable

Remote

AI Researcher
No salary listedRemoteAI Research

Conduct applied research on AI agents, designing experiments and evaluation systems to improve reliability, context retention, and multi-step task completion. The role requires strong AI/ML research, engineering, experimental design, and communication skills.