Machine Learning Data Engineer
Own the data foundation for machine-learning systems by building dataset pipelines, labeling workflows, quality controls, and lineage processes. The role requires at least three years of data engineering experience, strong Python and SQL skills, and familiarity with ML data quality concerns and orchestration tools.
About the job
Responsibilities
- Build and maintain dataset ingestion, cleaning, and versioning pipelines for training, validation, and held-out evaluation sets.
- Design and operate labeling workflows, including quality control and inter-annotator agreement tracking.
- Own dataset testing, including coverage analysis, class-balance checks, leakage detection, and dataset documentation.
- Manage dataset versioning and lineage so every model release traces back to its exact training data.
- Partner with data science and product engineering teams on data-sharing boundaries and formats.
Requirements
- 3+ years of data engineering experience.
- Strong Python and SQL skills.
- Experience with data versioning tools such as DVC or LakeFS, or similar tools.
- Experience with pipeline orchestration tools such as Airflow or Dagster, or similar tools.
- Understanding of machine-learning data concerns, including train/test leakage, label noise, and distribution shift.
- Rigor in documentation and reproducibility.
Benefits
- Innovative cybersecurity environment working with cutting-edge technologies.
- Collaborative culture that values team input.
- Ongoing training and professional development opportunities.
- Mission-driven work protecting critical infrastructure and digital assets.
- Team-building activities and social events.
Skills
Python, SQL, Dvc, Lakefs, Airflow, Dagster, Data Versioning, Pipeline Orchestration, Dataset Versioning, Data Labeling, Data Quality, Machine Learning
Similar jobs
Data Engineering jobsOwn the systems that ingest, standardize, validate, and operationalize data signals for Vanta’s EPD organization. The role suits a hands-on builder who has recently shipped working tools or pipelines, uses AI-assisted development, and helps teammates grow technically.
Build and operate scalable lakehouse infrastructure, streaming and CDC pipelines, query systems, and self-serve BI capabilities. Requires 5+ years of data engineering experience, strong Kubernetes and infrastructure-as-code expertise, and hands-on experience with distributed data platforms.
The Senior Platform Engineer will build and operate reliable data platform tooling, consolidate orchestration, scale dbt infrastructure, and improve Databricks developer experience. The role requires 5+ years of production software experience, strong Python and AWS expertise, infrastructure-as-code experience, and familiarity with modern data stacks.
Senior Data Engineer responsible for designing and deploying scalable data infrastructure, orchestration models, and analytics tooling to enable data-driven decisions, ML products, and enterprise reporting at Vanta. Requires 4+ years data experience, software engineering mindset, modern data stack proficiency, and passion for secure, compliant data systems.
Senior Analytics Engineer responsible for designing complex data models, building scalable SQL pipelines, enabling AI tooling, and improving data infrastructure to support self-serve analytics, dashboards, and data science at Vanta. Requires 4+ years data experience, software engineering mindset, and expertise with modern analytics tools like dbt.