Research Scientist
Conducts foundational and post-training research for large-scale video generation models while building scalable multimodal data pipelines. The role requires hands-on experience with distributed ML systems, distillation, reward modeling, preference-based fine-tuning, and Python-based deep learning frameworks.
About the job
Responsibilities
- Build and maintain scalable systems for ingesting, preprocessing, and delivering large-scale video data for model training.
- Design and scale distributed data pipelines for preprocessing, dataset generation, and repeated dataset refreshes.
- Own workflow orchestration, job scheduling, monitoring, and failure recovery for large-scale data processing jobs.
- Implement and maintain containerized pipeline infrastructure using Kubernetes or equivalent orchestration systems.
- Optimize cloud-based data storage and movement across AWS, Google Cloud Storage, or Azure for cost, throughput, and operational efficiency.
- Define and implement best practices for dataset storage layout, versioning, caching, retention, and access patterns.
- Build tooling to support deduplication workflows at scale, including near-deduplication pipelines over large video corpora.
- Research and develop distillation methods for large-scale diffusion and flow-based video generation models, including guidance and adversarial distillation.
- Develop reward models and preference-based fine-tuning pipelines aligned with human judgments of aesthetics, motion quality, and prompt adherence.
- Analyze the relationship between base model behavior and post-training outcomes, and inform pretraining decisions with the foundation model team.
Requirements
- Strong hands-on experience building or scaling large-scale data systems or machine learning pipelines.
- Experience with distributed data processing frameworks such as PySpark or Ray, and orchestration tools such as Airflow or equivalent.
- Familiarity with containerization and orchestration, including Docker and Kubernetes.
- Experience with cloud-based data storage and compute using AWS, GCS, and/or Azure, including cost, throughput, storage layout, and access-pattern tradeoffs.
- Familiarity with video and media processing tools such as FFmpeg, PyAV, DALI, or OpenCV.
- Familiarity with multimodal or media data, including video, image, text, and audio.
- Strong research background in post-training methods for large-scale diffusion or flow-based generative models, with hands-on experience in distillation for inference efficiency and quality preservation.
- Experience with reward modeling or preference-based fine-tuning for generative models, including RLHF, DPO, or equivalent alignment approaches.
- Solid understanding of the interplay between pretraining and post-training, including how base-model properties affect distillation and fine-tuning outcomes.
- Proficiency in Python and modern machine learning frameworks, with a strong preference for PyTorch or JAX.
- Track record of independent research and the ability to drive projects from initial idea through experimental validation.
- Good understanding of the practical challenges involved in building reliable, scalable, and reproducible machine learning data workflows.
Nice to Have
- Publications at top-tier venues such as NeurIPS, ICML, ICLR, CVPR, ICCV, or ECCV.
Benefits and Compensation
- Competitive salary and generous company equity.
- Personal time off and paid holidays.
- Health insurance.
- Global travel insurance covering international travel.
- Monthly spending stipend of $500 (~S$635).
- Equipment for a home office.
Skills
Python, PyTorch, JAX, Pyspark, Ray, Airflow, Docker, Kubernetes, AWS, GCP, Azure, Ffmpeg, Pyav, Opencv, Diffusion Models
Similar jobs
ML Engineering jobsBuild the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.
Build production agent systems that plan, use tools, recover from failures, and improve over time. The role requires 5+ years of production ML or backend experience, LLM or agent deployment experience, and expertise in evaluation, tracing, observability, and agent architecture.
Own inference-stack cost and performance by optimizing serving, caching, batching, quantization, decoding, routing, and GPU execution. The role requires 5+ years in ML systems, inference infrastructure, or performance engineering, plus strong Python and systems-language skills.
Conduct research and engineering on next-generation video generation models, including post-training, evaluation, multimodal data systems, and scalable ML infrastructure. The internship requires current bachelor's or master's study and practical experience with Python, data pipelines, or distributed systems.
Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.