Member of Technical Staff — ML Data Infra
Build and operate large-scale multimodal data pipelines for AI avatar model training. Design production-grade systems for petabyte-scale video, audio, and text data.
About the job
What You'll Do
- Design, build, and operate large-scale data pipelines for ingestion, processing, filtering, and curation of multimodal training data (video, audio, text)
- Take research-grade data processing code and turn it into robust, production-level pipelines — quickly and without losing correctness
- Optimize pipeline throughput and efficiency at scale; identify and eliminate bottlenecks across compute, I/O, and storage
- Build and maintain data quality systems — deduplication, filtering, validation, and quality scoring at scale
- Manage petabyte-scale datasets: storage architecture, versioning, lineage tracking, and cost efficiency
- Work closely with researchers to understand data requirements and translate them into scalable processing systems
- Build tooling and infrastructure that makes the research team faster — efficient data access, reproducible processing, and fast iteration loops
What We're Looking For
- Proven experience building and operating large-scale data pipelines in production — you've processed data at a scale where naive approaches break
- Strong proficiency with distributed data processing frameworks — Spark, Ray, Dask, or similar — and a clear sense of when to use each
- Solid software engineering fundamentals: you write clean, testable, maintainable code and understand why that matters when pipelines run unattended at scale
- Experience with multimodal data (video, audio) is a strong plus — understanding of formats, codecs, and processing libraries (FFmpeg, decord, etc.)
- Familiarity with ML data pipelines specifically — understanding of how data quality and format affect model training
- Ability to move fast: you can take a prototype script from a researcher and ship a production version in days, not weeks
Bonus Points
- Experience building data pipelines for large-scale model training (pre-training or fine-tuning)
- Familiarity with data versioning and lineage tools (DVC, Delta Lake, Apache Iceberg, etc.)
- Experience with streaming data pipelines or online data processing
- Prior work at an AI lab, video platform, or other data-intensive company
- Contributions to open-source data tooling
Compensation
- $200,000 – $300,000 base salary, plus meaningful equity
- Health: HSA plan with ~$2,000 in company contributions
- PTO: 15 days + public holidays, and we close for a full week over the holidays
- Lunch, beverages, and snacks on us every workday
- Commuter benefits
- 401K: In the works
Skills
Spark, Ray, Dask, Ffmpeg, Dvc, Delta Lake, Apache Iceberg, Python, Distributed Data Processing, Multimodal Data Processing
Similar jobs
Data Engineering jobsBuilds scalable data pipelines and data engine architecture for machine learning, integrating foundation models to automate labeling and discovery. The role requires 5+ years of experience, modern ML infrastructure expertise, and U.S. citizenship with security-clearance eligibility.
Builds and optimizes scalable data pipelines, storage, and OLAP databases for ML training, analytics, and product features. Requires 5+ years in data engineering, proficiency in Python/SQL/cloud platforms, and distributed systems experience.
Build and optimize Ray Data, a Python-native data processing engine for large-scale AI workloads. The role focuses on distributed systems performance, scalable data pipelines, production training solutions, and fault tolerance while partnering with AI-focused customers.
Build scalable data pipelines, infrastructure, and quantitative models that support experimentation, forecasting, and business decision-making. The role requires 4+ years of production data engineering experience, strong Python and SQL skills, distributed computing expertise, and a quantitative degree.
Own the systems that ingest, standardize, validate, and operationalize data signals for Vanta’s EPD organization. The role suits a hands-on builder who has recently shipped working tools or pipelines, uses AI-assisted development, and helps teammates grow technically.