Data Engineer, Machine Learning
Build and maintain production data pipelines that prepare conversational, voice, and multimodal data for ML model training and evaluation. Partner closely with ML engineers to deliver high-quality, versioned datasets and infrastructure.
About the job
Responsibilities
- Design and build production data pipelines that prepare conversational, voice, and multimodal data for model training and evaluation.
- Partner directly with ML engineers to understand data requirements for new models and experiments, and deliver datasets that meet those needs.
- Build and maintain infrastructure for dataset versioning, lineage tracking, and reproducibility.
- Develop data quality frameworks: schema validation, drift detection, and coverage monitoring.
- Optimise large-scale data processing for cost and performance across cloud infrastructure.
- Build tooling that makes it easy for ML engineers and researchers to discover, explore, and request data independently.
- Define and enforce data governance and privacy standards, particularly around sensitive conversational and voice data.
- Contribute to architecture decisions around the broader data platform.
Requirements
- 5+ years in data engineering, with meaningful experience supporting ML or AI teams.
- Strong SQL and Python skills.
- Experience building and operating ETL/ELT pipelines at scale using modern data platforms and tooling.
- Experience with workflow orchestration systems such as Airflow, Dagster, or Prefect.
- Hands-on experience with ML data workflows: training data pipelines, dataset versioning, data labeling pipelines, or model evaluation data.
- Solid understanding of how ML teams work and what makes a good training dataset.
- Comfort working with unstructured and semi-structured data — audio, text, JSON logs.
- Strong communication skills.
Nice-to-Haves
- Vector databases, embedding storage, or feature stores.
- Data from hardware or embedded systems: telemetry, sensors, real-time streams.
- Distributed compute frameworks such as Ray or Spark.
- Kubernetes and managed Kubernetes environments such as GKE or EKS.
- Data privacy frameworks, especially around voice or conversational data.
- Building internal tooling or self-serve data platforms.
Benefits
- 401(k) max employer match: 3.5% of compensation
- 100% employer-paid health, vision, and dental benefits for you and your dependents
- Unlimited PTO and sick time
- Flexible spending account with employer matching up to $1,650/year (medical FSA)
- Guardian Employee Assistance Program (EAP)
- Competitive stock options
Skills
Python, SQL, Airflow, Dagster, Prefect, ETL, ELT, Dataset Versioning, Data Quality, Ray, Spark, Kubernetes, GKE, EKS, Vector Databases
Similar jobs
Data Engineering jobsBuild and operate scalable data infrastructure, including partner data sharing, identity graph foundations, and governed batch and real-time platforms. The role requires 5+ years of data, distributed systems, infrastructure, or backend engineering experience and strong cloud and data-platform expertise.
Own the systems that ingest, standardize, validate, and operationalize data signals for Vanta’s EPD organization. The role suits a hands-on builder who has recently shipped working tools or pipelines, uses AI-assisted development, and helps teammates grow technically.
Build scalable data pipelines, infrastructure, and quantitative models that support experimentation, forecasting, and business decision-making. The role requires 4+ years of production data engineering experience, strong Python and SQL skills, distributed computing expertise, and a quantitative degree.
Builds and optimizes scalable data pipelines, storage, and OLAP databases for ML training, analytics, and product features. Requires 5+ years in data engineering, proficiency in Python/SQL/cloud platforms, and distributed systems experience.
Build and operate reliable, production-grade data pipelines, warehouse infrastructure, and trusted datasets supporting company-wide analytics and AI initiatives. The role requires 3+ years of production data engineering experience, strong SQL and Python skills, and experience with Snowflake, dbt, cloud infrastructure, and orchestration.