Machine Learning Engineer
Build and scale distributed pipelines that ingest, curate, filter, and prepare large-scale video and multimodal datasets for model training. The role requires Python, distributed processing, orchestration, containers, cloud infrastructure, and experience with VLM-based captioning or data-quality workflows.
About the job
Responsibilities
- Design and scale distributed data pipelines for preprocessing, dataset generation, and repeated dataset refreshes.
- Own workflow orchestration, job scheduling, monitoring, and failure recovery for large-scale data processing jobs.
- Implement and maintain containerized pipeline infrastructure using Kubernetes or equivalent orchestration systems.
- Optimize cloud-based data storage and movement across providers for cost, throughput, and operational efficiency.
- Define and implement best practices for dataset storage layout, versioning, caching, retention, and access patterns.
- Design and implement curation pipelines to select, filter, and retain video and image content for model training, including image-text pair datasets for joint training regimes.
- Build and improve VLM-based captioning and metadata-generation workflows at scale across video and image data.
- Develop and apply quality and aesthetic scoring models, CLIP-based semantic filtering, and other signal-extraction approaches for data selection.
- Build tooling for large-scale deduplication workflows, including near-deduplication and exact deduplication over large video corpora.
- Analyze dataset composition, identify quality issues, and iterate on curation logic to improve training outcomes.
- Define and evolve standards for high-quality, training-ready video data across different training regimes.
Requirements
- Strong hands-on experience building or scaling large-scale machine-learning data systems and pipelines, including dataset curation, filtering, and quality improvement.
- Experience with distributed data-processing frameworks such as PySpark or Ray, and workflow orchestration tools such as Airflow or equivalent.
- Familiarity with containerization and container orchestration, including Docker and Kubernetes.
- Experience with cloud-based data storage and compute, including AWS, GCS, and/or Azure, and tradeoffs involving cost, throughput, storage layout, and access patterns.
- Experience with VLM-based captioning pipelines or quality/aesthetic scoring models for video or image data, including curation of image-text pair datasets for joint image-video training.
- Familiarity with CLIP-based or embedding-based filtering and semantic data-selection techniques.
- Familiarity with video and media-processing tools such as FFmpeg, PyAV, DALI, or OpenCV, and libraries such as Decord, torchvision, PyTorchVideo, or torchaudio.
- Proficiency in Python.
- Strong problem-solving, communication, and documentation skills.
Benefits
- Competitive salary and generous company equity.
- Personal time off and paid holidays.
- Health insurance.
- Global travel insurance covering international travel.
- Monthly spending stipend of $500 (~S$635).
- Equipment provided for a home office.
Skills
Python, Pyspark, Ray, Airflow, Docker, Kubernetes, AWS, GCP, Azure, Vlms, Clip, Ffmpeg, Pyav, Opencv, PyTorch
Similar jobs
ML Engineering jobsBuild the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.
Build production agent systems that plan, use tools, recover from failures, and improve over time. The role requires 5+ years of production ML or backend experience, LLM or agent deployment experience, and expertise in evaluation, tracing, observability, and agent architecture.
Own inference-stack cost and performance by optimizing serving, caching, batching, quantization, decoding, routing, and GPU execution. The role requires 5+ years in ML systems, inference infrastructure, or performance engineering, plus strong Python and systems-language skills.
Conduct research and engineering on next-generation video generation models, including post-training, evaluation, multimodal data systems, and scalable ML infrastructure. The internship requires current bachelor's or master's study and practical experience with Python, data pipelines, or distributed systems.
Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.