Member of Technical Staff - Post Training
Owns end-to-end post-training for frontier multimodal generative models, spanning reward modeling, preference optimization, distillation, safety tuning, evaluation, and deployment. The role requires prior experience shipping post-training improvements and strong PyTorch expertise.
About the job
Responsibilities
- Own the full post-training pipeline end to end, from data curation and reward modeling through fine-tuning, preference optimization, distillation, safety tuning, evaluation, and deployment.
- Advance post-training techniques including SFT, RLHF, RLAIF, DPO, preference learning, and reward modeling to align models with human intent and aesthetic judgment.
- Work across text-to-image, image editing, multi-reference, and video post-training.
- Build personalization and customization capabilities that let users adapt models to their creative style.
- Design and maintain high-throughput fine-tuning and evaluation infrastructure for rapid research iteration.
- Identify quality and alignment gaps through rigorous evaluation, then address them through targeted research and engineering.
Requirements
- Experience owning post-training for a frontier generative model through release, including SFT, preference optimization, distillation, and safety tuning, with measurable quality gains on human preferences or standard benchmarks.
- Deep experience across reward modeling, preference learning, RLHF/RLAIF, and personalization.
- Experience working across modalities, including text-to-image, image editing, and multi-reference; video experience is preferred.
- Strong PyTorch fluency and ability to write maintainable research code.
- Bias toward shipping measurable model-quality improvements that reach users.
Nice to Have
- Experience with distillation techniques such as LADD, DMD, or consistency models.
- Experience building high-throughput evaluation pipelines.
Compensation
- Base annual salary: €130,000–€340,000, plus equity.
Skills
PyTorch, Sft, RLHF, Rlaif, Dpo, Reward Modeling, Preference Learning, Distillation, Safety Tuning, Generative Models, Multimodal Models, Evaluation Pipelines, Ladd, Dmd, Consistency Models
Similar jobs
ML Engineering jobsBuild and operate AI-powered, customer-facing workflows for Datadog Notebooks, combining reliable backend systems with LLM capabilities. The role requires 6+ years of engineering experience, Go or Python expertise, and experience delivering production AI products.
Trains and fine-tunes large-scale diffusion transformer models for image and video generation, conducts rigorous ablation studies, and optimizes distributed training. Requires hands-on diffusion-model experience, strong PyTorch and transformer expertise, and understanding of generative-model evaluation.
Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.
Build the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.
Build production agent systems that plan, use tools, recover from failures, and improve over time. The role requires 5+ years of production ML or backend experience, LLM or agent deployment experience, and expertise in evaluation, tracing, observability, and agent architecture.