Latest ML Engineering jobs at Cantina
Job results
Conduct research and engineering on next-generation video generation models, including post-training, evaluation, multimodal data systems, and scalable ML infrastructure. The internship requires current bachelor's or master's study and practical experience with Python, data pipelines, or distributed systems.
Build and productionize large-scale generative speech models for voice conversion and related capabilities. The role combines research, data, evaluation, distributed training, performance optimization, and responsible deployment of speech systems.
Build state-of-the-art end-to-end speech and audio generation systems with a focus on joint audio-video modeling. Own audio representations (VAEs, neural codecs), generative backbones (diffusion/flow-matching transformers), conditioning, alignment for voice cloning and sync with video, plus data flywheel, evaluation, and inference optimization for large-scale multimodal models.
Build and scale inference infrastructure for generative audio models including TTS, voice conversion, and ASR. Design high-performance, low-latency serving systems using Kubernetes, CI/CD, and GPU optimization to bridge research and production.
Build and scale distributed pipelines that ingest, curate, filter, and prepare large-scale video and multimodal datasets for model training. The role requires Python, distributed processing, orchestration, containers, cloud infrastructure, and experience with VLM-based captioning or data-quality workflows.
Conducts foundational and post-training research for large-scale video generation models while building scalable multimodal data pipelines. The role requires hands-on experience with distributed ML systems, distillation, reward modeling, preference-based fine-tuning, and Python-based deep learning frameworks.
Designs evaluation pipelines, metrics, and user studies for speech generation and recognition models (ASR/TTS). Trains evaluation models, builds dashboards, and collaborates with ML, data, and product teams to improve performance on large-scale systems.
Designs, fine-tunes, and deploys image generation models for photorealistic AI bots, optimizing for consistency, latency, and quality. Requires 5+ years software engineering, 2+ years production ML, and expertise in diffusion models like Stable Diffusion and PyTorch.