# Machine Learning Engineer, Speech

**Company:** [Cantina](https://hotfix.jobs/companies/cantina)
**Location:** Remote
**Role:** ML Engineering
**Salary:** $200k – $220k/yr
**Experience:** 7+ years
**Skills:** PyTorch, diffusion models, flow matching, audio vaes, neural audio codecs, vocoders, fsdp, deepspeed, CUDA, triton, voice cloning, multimodal modeling, video generation, Distributed Training
**Posted:** 2026-07-30

> Build state-of-the-art end-to-end speech and audio generation systems with a focus on joint audio-video modeling. Own audio representations (VAEs, neural codecs), generative backbones (diffusion/flow-matching transformers), conditioning, alignment for voice cloning and sync with video, plus data flywheel, evaluation, and inference optimization for large-scale multimodal models.

## Job Description

## What You’ll Do

**Audio Representations:** Design, train, and improve the audio VAEs, neural codecs, and vocoders our generative models sit on top of latent design, reconstruction and perceptual objectives, compression-vs-fidelity tradeoffs.

**Model Building:** Architect, implement, pre-train, fine-tune, and post-train/alignment (e.g., GRPO/DPO) diffusion and flow-matching transformers for large-scale audio and video generation.

**Joint Audio-Video Modeling:** Design the audio conditioning and cross-modal alignment inside joint AV models, audio latents alongside video latents, reference-audio and multi-speaker conditioning, multi shot generation audio/video modeling.

**Experimental Design:** Design, run, and analyze scientific experiments to advance our understanding of the models.

**Data Ownership:** Define data requirements and collaborate on acquisition, curation, AV-sync and quality filtering, annotation quality, and synthetic data strategies for paired audio-video and speech corpora.

**Rigorous Evaluation:** Design automated objective/subjective evaluations audio fidelity and intelligibility metrics, AV-sync, listening and viewing tests, robustness & bias checks, and red-team studies.

**Inference Efficiency:** Drive distillation, step-count reduction, quantization, and kernel/memory optimization to meet interactive latency and cost targets.

**Pipeline Delivery:** Harden the training → evaluation → inference pipeline; profile latency, memory, and cost; and meet production SLAs with robust monitoring and rollback.

**GPU Scaling:** Partner with infrastructure to run distributed training/inference on cloud fleets and productionize models with reliability and observability.

**Project Leadership:** Independently lead small research projects while collaborating on larger team initiatives, including cross-team work with video generation.

**Tool Development:** Develop and improve dev tooling to enhance team productivity.

**Safety & Responsibility:** Contribute to safety/consent guardrails, watermarking, and misuse/abuse mitigation for responsible voice and likeness technology.

## What You’ll Bring

- Exceptional research/development experience with large-scale audio models (>8B parameters, >500k hours of data).
- Deep hands-on experience with diffusion and/or flow-matching transformers, including practical knowledge of samplers, schedules, conditioning mechanisms, and distillation.
- Deep hands-on experience training audio VAEs, neural audio codecs, and vocoders latent/tokenizer design, reconstruction and perceptual objectives, adversarial training.
- Strong experience with multi-node, multi-GPU distributed training (FSDP/DeepSpeed or equivalent).
- Strong software engineering skills with a proven track record of building complex systems.
- Strong with PyTorch and performance work (profiling, CUDA/Triton/C++ as needed) and writing reliable production-quality code.
- Shipped large-scale speech/audio or multimodal generative models to production.
- Background in working with large-scale ML data, and the ability to iterate on data and triangulate quality using both subjective and objective signals.
- Experience with voice cloning, speech control/steerability, or expressive speech generation.
- Notable publications and/or open-source contributions in speech/audio/ML.

**Strongly preferred:**
- Experience with multimodal audio-video modeling: joint AV generation of multi-shot, multi-speaker scenes with dialogue, music, and sound design generated jointly with video, and the cross-modal alignment that keeps them in sync.
- Experience with video generation: video diffusion/flow-matching transformers, video VAEs, conditioned and multi-shot generation, building data pipelines for video models.
- Streaming or real-time generation, causal distillation (e.g., Self Forcing / Self Forcing++).

## Compensation

The anticipated annual base salary range for this role is between $200,000-$220,000 (€170,000-€190,000). When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.

## Benefits for U.S.-based roles

- Competitive salary and generous company equity
- Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina
- 42 days of paid time off, including: 15 PTO days, 10 sick days, 15 company holidays, 2 floating holidays
- Generous parental leave & fertility support
- 401(k) retirement savings plan
- Lifestyle spending account – $500/month to use however you’d like
- Complimentary lunch and snacks for in-office employees
- One Medical membership, and more!

## Similar roles

- [Senior Backend Engineer](https://hotfix.jobs/jobs/6915e8e1-b4db-4143-8f8f-24d2a075f6a4) - Tennr - New York, NY - $200k – $230k/yr
- [Senior Software Engineer, Build](https://hotfix.jobs/jobs/a127ac15-4622-4734-8aa4-0d955d917a99) - Astronomer - New York, NY - $200k – $230k/yr
- [Senior Software Engineer, Backend](https://hotfix.jobs/jobs/f97d2d53-1f1a-4ae9-9d09-2a05adeec807) - Siftstack - Marina Del Rey, CA - $200k – $250k/yr
- [Senior Software Engineer](https://hotfix.jobs/jobs/17e0432e-a11f-4a57-8de2-0c603c7b5780) - Traba - New York, NY - $200k – $240k/yr
- [Machine Learning Research Manager](https://hotfix.jobs/jobs/d6047c03-2cd4-40fd-92a5-332a272510cc) - Rad AI - San Francisco, CA - $200k – $230k/yr

**Apply:** https://hotfix.jobs/jobs/eaf1d394-7250-48b6-9d9d-98ff4ceeca1f
**Canonical:** https://hotfix.jobs/jobs/eaf1d394-7250-48b6-9d9d-98ff4ceeca1f