Machine Learning Intern
Conduct research and engineering on next-generation video generation models, including post-training, evaluation, multimodal data systems, and scalable ML infrastructure. The internship requires current bachelor's or master's study and practical experience with Python, data pipelines, or distributed systems.
About the job
Responsibilities
- Develop reward models to improve aesthetics, motion quality, temporal consistency, and prompt adherence in video generation models.
- Study base-model behavior and use experimental findings to inform model development.
- Design evaluations and conduct large-scale experiments on generative video models.
- Build systems for ingesting, preprocessing, curating, and delivering large-scale video datasets.
- Develop distributed pipelines for dataset generation, deduplication, preprocessing, and recurring dataset refreshes.
- Improve the reliability, reproducibility, and efficiency of data and model-training workflows.
- Build tooling for video and multimodal data using FFmpeg, PyAV, DALI, OpenCV, or equivalent technologies.
- Contribute to evaluation harnesses, model integrations, research tooling, and related product-adjacent projects.
- Document and communicate findings through research reports, presentations, demonstrations, and potential conference submissions.
Requirements
- Pursuing a bachelor's or master's degree in computer science, engineering, machine learning, or a related field.
- Experience building data pipelines, ML infrastructure, or distributed systems through research, coursework, open-source work, or internships.
- Familiarity with Ray, PySpark, Airflow, Docker, Kubernetes, or equivalent technologies.
- Experience with cloud storage or compute platforms such as AWS, Google Cloud, or Azure.
- Understanding of data throughput, storage layout, caching, monitoring, and failure recovery.
- Proficiency in Python and interest in building reliable systems for large-scale machine learning.
Nice to Have
- Experience with video, image, audio, or other multimodal data.
- Publications at venues such as NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV, or AAAI.
Compensation and Benefits
- Competitive monthly stipend.
- Visa and travel support for eligible international candidates.
- Housing support for qualifying international interns in Singapore.
- Defined project, named senior mentor, weekly one-on-one meetings, and project milestones.
- Compute allocation, equipment, and resources for research.
- Conference travel support for accepted papers.
- Three-month onsite internship with batches starting in October 2026 and January 2027.
Skills
Python, Ray, Pyspark, Airflow, Docker, Kubernetes, AWS, GCP, Azure, Ffmpeg, Pyav, Dali, Opencv, Distributed Systems, ML Infrastructure
Similar jobs
ML Engineering jobsBuild the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.
Build production agent systems that plan, use tools, recover from failures, and improve over time. The role requires 5+ years of production ML or backend experience, LLM or agent deployment experience, and expertise in evaluation, tracing, observability, and agent architecture.
Own inference-stack cost and performance by optimizing serving, caching, batching, quantization, decoding, routing, and GPU execution. The role requires 5+ years in ML systems, inference infrastructure, or performance engineering, plus strong Python and systems-language skills.
Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.