# Research Scientist

**Company:** [Cantina](https://hotfix.jobs/companies/cantina)
**Location:** Unspecified
**Role:** ML Engineering
**Skills:** Python, PyTorch, JAX, Pyspark, Ray, Airflow, Docker, Kubernetes, AWS, GCP, Azure, Ffmpeg, Pyav, Opencv, Diffusion Models
**Posted:** 2026-05-12

> Conducts foundational and post-training research for large-scale video generation models while building scalable multimodal data pipelines. The role requires hands-on experience with distributed ML systems, distillation, reward modeling, preference-based fine-tuning, and Python-based deep learning frameworks.

## Job Description

## Responsibilities
- Build and maintain scalable systems for ingesting, preprocessing, and delivering large-scale video data for model training.
- Design and scale distributed data pipelines for preprocessing, dataset generation, and repeated dataset refreshes.
- Own workflow orchestration, job scheduling, monitoring, and failure recovery for large-scale data processing jobs.
- Implement and maintain containerized pipeline infrastructure using Kubernetes or equivalent orchestration systems.
- Optimize cloud-based data storage and movement across AWS, Google Cloud Storage, or Azure for cost, throughput, and operational efficiency.
- Define and implement best practices for dataset storage layout, versioning, caching, retention, and access patterns.
- Build tooling to support deduplication workflows at scale, including near-deduplication pipelines over large video corpora.
- Research and develop distillation methods for large-scale diffusion and flow-based video generation models, including guidance and adversarial distillation.
- Develop reward models and preference-based fine-tuning pipelines aligned with human judgments of aesthetics, motion quality, and prompt adherence.
- Analyze the relationship between base model behavior and post-training outcomes, and inform pretraining decisions with the foundation model team.

## Requirements
- Strong hands-on experience building or scaling large-scale data systems or machine learning pipelines.
- Experience with distributed data processing frameworks such as PySpark or Ray, and orchestration tools such as Airflow or equivalent.
- Familiarity with containerization and orchestration, including Docker and Kubernetes.
- Experience with cloud-based data storage and compute using AWS, GCS, and/or Azure, including cost, throughput, storage layout, and access-pattern tradeoffs.
- Familiarity with video and media processing tools such as FFmpeg, PyAV, DALI, or OpenCV.
- Familiarity with multimodal or media data, including video, image, text, and audio.
- Strong research background in post-training methods for large-scale diffusion or flow-based generative models, with hands-on experience in distillation for inference efficiency and quality preservation.
- Experience with reward modeling or preference-based fine-tuning for generative models, including RLHF, DPO, or equivalent alignment approaches.
- Solid understanding of the interplay between pretraining and post-training, including how base-model properties affect distillation and fine-tuning outcomes.
- Proficiency in Python and modern machine learning frameworks, with a strong preference for PyTorch or JAX.
- Track record of independent research and the ability to drive projects from initial idea through experimental validation.
- Good understanding of the practical challenges involved in building reliable, scalable, and reproducible machine learning data workflows.

## Nice to Have
- Publications at top-tier venues such as NeurIPS, ICML, ICLR, CVPR, ICCV, or ECCV.

## Benefits and Compensation
- Competitive salary and generous company equity.
- Personal time off and paid holidays.
- Health insurance.
- Global travel insurance covering international travel.
- Monthly spending stipend of $500 (~S$635).
- Equipment for a home office.

## Similar jobs

- [AI Engineer - New Verticals](https://hotfix.jobs/jobs/4159c72e-536e-4211-969c-6bcb9ad805fd) - Protege - Remote
- [Agent Systems Engineer](https://hotfix.jobs/jobs/2b04a068-018b-4fe0-915f-4fdf80d4c582) - Adaption Labs - San Francisco, CA
- [Inference Performance Engineer](https://hotfix.jobs/jobs/7718c69a-6d49-445e-8c7c-840bd1b3f346) - Adaption Labs - San Francisco, CA
- [Machine Learning Intern](https://hotfix.jobs/jobs/6c35fd46-271b-4839-88c3-44ec358c478d) - Cantina
- [Staff Machine Learning Engineer](https://hotfix.jobs/jobs/2cb6da8f-4a54-448d-a4b6-fc05f1412a12) - Payabli - Remote

**Apply:** https://hotfix.jobs/jobs/a51b7963-6b0f-411c-9f1d-ec1cde2e54e7
**Canonical:** https://hotfix.jobs/jobs/a51b7963-6b0f-411c-9f1d-ec1cde2e54e7