Skip to content
CantinaCantinaUnited States

Machine Learning Engineer - Voice Conversion

Build and productionize large-scale generative speech models for voice conversion and related capabilities. The role combines research, data, evaluation, distributed training, performance optimization, and responsible deployment of speech systems.

200k – 220k/yr
RemoteML Engineering

About the role

Responsibilities

  • Architect, implement, pre-train, fine-tune, and post-train or align large-scale speech models, including GRPO and DPO approaches.
  • Design, run, and analyze scientific experiments to advance model understanding.
  • Develop and improve developer tooling to enhance team productivity.
  • Contribute across the stack, from low-level optimizations to high-level model design.
  • Define data requirements and collaborate on acquisition, curation, augmentation, labeling quality, and synthetic data strategies.
  • Design automated objective and subjective evaluations, including listening tests, SV/WER/ASR-based metrics, robustness and bias checks, and red-team studies.
  • Harden training, evaluation, and inference pipelines; profile latency, memory, and cost; and meet production SLAs with monitoring and rollback.
  • Contribute to safety and consent guardrails and misuse or abuse mitigation for speech technology.

Requirements

  • Exceptional research or development experience with large-scale audio models exceeding 8B parameters and 500,000 hours of data.
  • Deep hands-on experience with diffusion and/or flow-matching transformers, including samplers, schedules, conditioning mechanisms, and distillation.
  • Deep hands-on experience training audio VAEs, neural audio codecs, and vocoders, including latent/tokenizer design, reconstruction and perceptual objectives, and adversarial training.
  • Strong experience with multi-node, multi-GPU distributed training using FSDP, DeepSpeed, or equivalent technologies.
  • Strong software engineering skills and experience building complex systems.
  • Strong PyTorch and performance engineering skills, including profiling and use of CUDA, Triton, or C++ as needed.
  • Experience shipping large-scale speech, audio, or multimodal generative models to production.
  • Experience working with large-scale ML data and evaluating quality using subjective and objective signals.
  • Experience with voice cloning, speech control or steerability, or expressive speech generation.
  • Notable publications and/or open-source contributions in speech, audio, or machine learning.

Compensation and Benefits

  • Annual base salary: $200,000-$220,000 (€170,000-€190,000).
  • Competitive salary and generous company equity.
  • Medical, dental, and vision insurance, with 99.99% of premiums covered by Cantina.
  • 42 days of paid time off, including 15 PTO days, 10 sick days, 15 company holidays, and 2 floating holidays.
  • Generous parental leave and fertility support.
  • 401(k) retirement savings plan.
  • $500/month lifestyle spending account.
  • Complimentary lunch and snacks for in-office employees.
  • One Medical membership.

Skills

PyTorchdiffusion modelsflow matchingTransformersaudio vaesneural audio codecsvocodersfsdpdeepspeedCUDAtritonC++voice cloningDistributed Traininggrpo

Similar roles

ML Engineering jobs
Together AI

Research Engineer, Large-Scale Training

Together AISan Francisco, CA

Research Engineer turning efficient foundation model training research into robust high-performance production systems at Together AI. Optimize large-scale training infrastructure, profile bottlenecks, integrate new models, and productionize novel methods in close partnership with scientists.

200k – 290k/yrOn-siteML Engineering
Siftstack

Software Engineer, Backend

SiftstackSan Francisco, CA +1

Build and operate dependable agentic AI systems that analyze large-scale hardware telemetry for aerospace and defense teams. Design tools, execution environments, distributed job systems on Kubernetes, and evaluation frameworks while owning the product end-to-end and speaking directly with customers.

200k – 250k/yrHybrid3+ YOEML Engineering
Console

Research Engineer

ConsoleSan Francisco, CA

Research Engineer building self-improving AI agent systems at Console. Develop eval/optimization loops, fine-tune specialist models, and improve agent reasoning over enterprise context using production data to drive measurable gains in quality, latency, and reliability.

200k – 350k/yrOn-siteML Engineering
Kepler

Machine Learning Engineer

KeplerNew York, NY

Build and own ML models, fine-tuning, evaluation harnesses, and routing for Kepler's AI agent harness in finance. Requires 5+ years production software experience and shipped ML systems focused on correctness, evals, and real-world reliability.

200k – 280k/yrOn-site5+ YOEML Engineering
AfterQuery

Machine Learning Engineer

AfterQuerySan Francisco, CA

Build production ML systems for measuring, predicting, and scaling data quality for frontier AI models. Requires 3-6 years experience in applied ML or related production systems (ranking, recommendations, data quality, fraud) plus strong software engineering skills.

200k – 300k/yrOn-site3+ YOEML Engineering