Skip to content
CantinaCantina

Member of Technical Staff, Data & ML Infrastructure for Video Models

Builds and scales data pipelines for video generation models, including ingestion, annotation via MTurk/Prolific, preprocessing, and curation using Python, AWS, Kubernetes. Requires 3+ years in ML/data engineering, PyTorch experience, and cross-functional collaboration.

About the job

Responsibilities

  • Build and maintain data pipelines for large video generation models, including data ingestion, parsing, filtering, preprocessing, and dataset curation at scale, using tools such as AWS S3 and DynamoDB.
  • Design and run annotation workflows across platforms such as MTurk, Prolific, including task design, quality control, and label validation.
  • Train, evaluate, and improve smaller supporting models used for data filtering, quality assessment, preprocessing, or other parts of the ML pipeline.
  • Partner closely with research and engineering teams to turn experimental workflows into scalable, repeatable systems that support model training and evaluation.
  • Own data quality across the pipeline by identifying bottlenecks, failure modes, and low-quality sources, and continuously improving tooling and processes.
  • Build internal tools and automation that make it easier to prepare datasets, launch annotation jobs, monitor outputs, and support model development end to end.
  • Drive larger pipeline projects from start to finish, such as new dataset creation efforts or upgrades to labeling and preprocessing infrastructure.
  • Work within a Kubernetes-based training infrastructure, ensuring datasets are properly prepared, formatted, and delivered to training clusters.
  • Profile and optimize research model inference scripts used in preprocessing steps, ensuring that model-driven filtering and transformation stages run within practical time and cost constraints when applied to large-scale raw data.

Requirements

  • 3+ years of experience in machine learning, applied ML, data pipelines, or related engineering roles, ideally working on large-scale multimodal, video, or vision-based systems.
  • Strong programming skills in Python and solid experience building reliable data processing and preprocessing pipelines for ML workflows.
  • Hands-on experience preparing training data for ML models, including parsing, filtering, dataset curation, quality control, and large-scale data handling using tools such as AWS S3 and DynamoDB.
  • Familiarity with annotation and labeling workflows, including task design, vendor or crowd-platform orchestration such as MTurk or Prolific, and methods for ensuring label quality.
  • Experience working with Kubernetes for orchestrating distributed workloads, including data preprocessing, pipeline execution, and dataset delivery to training clusters.
  • Comfort working across cloud and on-demand compute environments such as AWS and RunPod, with the ability to port and optimize pipelines across infrastructure.
  • Familiarity with distributed data processing frameworks and experience designing systems that operate reliably at scale across many nodes or workers.
  • Working knowledge of PyTorch and the broader deep learning stack, with the ability to read, debug, and optimize research model inference code for use in production preprocessing pipelines.
  • Ability to work cross-functionally with research and engineering teams and translate experimental ideas into robust, scalable systems.
  • Bachelor's, Master's, or PhD in Computer Science, Machine Learning, Engineering, Mathematics, or a related technical field; experience in generative video, computer vision, or multimodal ML is strongly preferred.

Nice-to-Haves

  • Experience training, evaluating, or fine-tuning smaller ML models used for classification, filtering, ranking, quality assessment, or other supporting tasks in an ML pipeline.

Compensation

  • Anticipated annual base salary range: $200,000-$260,000 (U.S.).

Skills

Python, Aws S3, DynamoDB, Kubernetes, PyTorch, Mturk, Prolific, Data Pipelines, Distributed Data Processing, Ml Preprocessing

Applied Intuition

Applied Intuition

Sunnyvale, CA

Data Engineer - Axion
$200k+/yrOn-site5+ YOEData Engineering

Builds scalable data pipelines and data engine architecture for machine learning, integrating foundation models to automate labeling and discovery. The role requires 5+ years of experience, modern ML infrastructure expertise, and U.S. citizenship with security-clearance eligibility.

Abridge

Abridge

San Francisco, CA

Data Engineer
$185k+/yrHybrid5+ YOEData Engineering

Builds and optimizes scalable data pipelines, storage, and OLAP databases for ML training, analytics, and product features. Requires 5+ years in data engineering, proficiency in Python/SQL/cloud platforms, and distributed systems experience.

Anyscale

Anyscale

San Francisco, CA

Software Engineer
$215k+/yrOn-site3+ YOEData Engineering

Build and optimize Ray Data, a Python-native data processing engine for large-scale AI workloads. The role focuses on distributed systems performance, scalable data pipelines, production training solutions, and fault tolerance while partnering with AI-focused customers.

xAI

xAI

Palo Alto, CA

Analytics Engineer - X
$180k+/yrOn-site4+ YOEData Engineering

Build scalable data pipelines, infrastructure, and quantitative models that support experimentation, forecasting, and business decision-making. The role requires 4+ years of production data engineering experience, strong Python and SQL skills, distributed computing expertise, and a quantitative degree.

Vanta

Vanta

Remote

Operations Manager, Signal Systems
$176k+/yrRemoteData Engineering

Own the systems that ingest, standardize, validate, and operationalize data signals for Vanta’s EPD organization. The role suits a hands-on builder who has recently shipped working tools or pipelines, uses AI-assisted development, and helps teammates grow technically.