Skip to content

Member of Technical Staff — ML Data Infra

Build and operate large-scale multimodal data pipelines for AI avatar model training. Design production-grade systems for petabyte-scale video, audio, and text data.

About the job

What You'll Do

  • Design, build, and operate large-scale data pipelines for ingestion, processing, filtering, and curation of multimodal training data (video, audio, text)
  • Take research-grade data processing code and turn it into robust, production-level pipelines — quickly and without losing correctness
  • Optimize pipeline throughput and efficiency at scale; identify and eliminate bottlenecks across compute, I/O, and storage
  • Build and maintain data quality systems — deduplication, filtering, validation, and quality scoring at scale
  • Manage petabyte-scale datasets: storage architecture, versioning, lineage tracking, and cost efficiency
  • Work closely with researchers to understand data requirements and translate them into scalable processing systems
  • Build tooling and infrastructure that makes the research team faster — efficient data access, reproducible processing, and fast iteration loops

What We're Looking For

  • Proven experience building and operating large-scale data pipelines in production — you've processed data at a scale where naive approaches break
  • Strong proficiency with distributed data processing frameworks — Spark, Ray, Dask, or similar — and a clear sense of when to use each
  • Solid software engineering fundamentals: you write clean, testable, maintainable code and understand why that matters when pipelines run unattended at scale
  • Experience with multimodal data (video, audio) is a strong plus — understanding of formats, codecs, and processing libraries (FFmpeg, decord, etc.)
  • Familiarity with ML data pipelines specifically — understanding of how data quality and format affect model training
  • Ability to move fast: you can take a prototype script from a researcher and ship a production version in days, not weeks

Bonus Points

  • Experience building data pipelines for large-scale model training (pre-training or fine-tuning)
  • Familiarity with data versioning and lineage tools (DVC, Delta Lake, Apache Iceberg, etc.)
  • Experience with streaming data pipelines or online data processing
  • Prior work at an AI lab, video platform, or other data-intensive company
  • Contributions to open-source data tooling

Compensation

  • $200,000 – $300,000 base salary, plus meaningful equity
  • Health: HSA plan with ~$2,000 in company contributions
  • PTO: 15 days + public holidays, and we close for a full week over the holidays
  • Lunch, beverages, and snacks on us every workday
  • Commuter benefits
  • 401K: In the works

Skills

Spark, Ray, Dask, Ffmpeg, Dvc, Delta Lake, Apache Iceberg, Python, Distributed Data Processing, Multimodal Data Processing

Applied Intuition

Applied Intuition

Sunnyvale, CA

Data Engineer - Axion
$200k+/yrOn-site5+ YOEData Engineering

Builds scalable data pipelines and data engine architecture for machine learning, integrating foundation models to automate labeling and discovery. The role requires 5+ years of experience, modern ML infrastructure expertise, and U.S. citizenship with security-clearance eligibility.

Abridge

Abridge

San Francisco, CA

Data Engineer
$185k+/yrHybrid5+ YOEData Engineering

Builds and optimizes scalable data pipelines, storage, and OLAP databases for ML training, analytics, and product features. Requires 5+ years in data engineering, proficiency in Python/SQL/cloud platforms, and distributed systems experience.

Anyscale

Anyscale

San Francisco, CA

Software Engineer
$215k+/yrOn-site3+ YOEData Engineering

Build and optimize Ray Data, a Python-native data processing engine for large-scale AI workloads. The role focuses on distributed systems performance, scalable data pipelines, production training solutions, and fault tolerance while partnering with AI-focused customers.

xAI

xAI

Palo Alto, CA

Analytics Engineer - X
$180k+/yrOn-site4+ YOEData Engineering

Build scalable data pipelines, infrastructure, and quantitative models that support experimentation, forecasting, and business decision-making. The role requires 4+ years of production data engineering experience, strong Python and SQL skills, distributed computing expertise, and a quantitative degree.

Vanta

Vanta

Remote

Operations Manager, Signal Systems
$176k+/yrRemoteData Engineering

Own the systems that ingest, standardize, validate, and operationalize data signals for Vanta’s EPD organization. The role suits a hands-on builder who has recently shipped working tools or pipelines, uses AI-assisted development, and helps teammates grow technically.