Machine Learning Engineer - Voice Conversion
Build and productionize large-scale generative speech models for voice conversion and related capabilities. The role combines research, data, evaluation, distributed training, performance optimization, and responsible deployment of speech systems.
About the job
Responsibilities
- Architect, implement, pre-train, fine-tune, and post-train or align large-scale speech models, including GRPO and DPO approaches.
- Design, run, and analyze scientific experiments to advance model understanding.
- Develop and improve developer tooling to enhance team productivity.
- Contribute across the stack, from low-level optimizations to high-level model design.
- Define data requirements and collaborate on acquisition, curation, augmentation, labeling quality, and synthetic data strategies.
- Design automated objective and subjective evaluations, including listening tests, SV/WER/ASR-based metrics, robustness and bias checks, and red-team studies.
- Harden training, evaluation, and inference pipelines; profile latency, memory, and cost; and meet production SLAs with monitoring and rollback.
- Contribute to safety and consent guardrails and misuse or abuse mitigation for speech technology.
Requirements
- Exceptional research or development experience with large-scale audio models exceeding 8B parameters and 500,000 hours of data.
- Deep hands-on experience with diffusion and/or flow-matching transformers, including samplers, schedules, conditioning mechanisms, and distillation.
- Deep hands-on experience training audio VAEs, neural audio codecs, and vocoders, including latent/tokenizer design, reconstruction and perceptual objectives, and adversarial training.
- Strong experience with multi-node, multi-GPU distributed training using FSDP, DeepSpeed, or equivalent technologies.
- Strong software engineering skills and experience building complex systems.
- Strong PyTorch and performance engineering skills, including profiling and use of CUDA, Triton, or C++ as needed.
- Experience shipping large-scale speech, audio, or multimodal generative models to production.
- Experience working with large-scale ML data and evaluating quality using subjective and objective signals.
- Experience with voice cloning, speech control or steerability, or expressive speech generation.
- Notable publications and/or open-source contributions in speech, audio, or machine learning.
Compensation and Benefits
- Annual base salary: $200,000-$220,000 (€170,000-€190,000).
- Competitive salary and generous company equity.
- Medical, dental, and vision insurance, with 99.99% of premiums covered by Cantina.
- 42 days of paid time off, including 15 PTO days, 10 sick days, 15 company holidays, and 2 floating holidays.
- Generous parental leave and fertility support.
- 401(k) retirement savings plan.
- $500/month lifestyle spending account.
- Complimentary lunch and snacks for in-office employees.
- One Medical membership.
Skills
PyTorch, Diffusion Models, Flow Matching, Transformers, Audio Vaes, Neural Audio Codecs, Vocoders, Fsdp, Deepspeed, CUDA, Triton, C++, Voice Cloning, Distributed Training, Grpo
Similar jobs
ML Engineering jobsBuild and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Build and operate production ML infrastructure spanning training, deployment, serving, monitoring, data pipelines, and feedback-driven retraining. The role requires strong MLOps and DevOps experience, Python and SQL proficiency, and ownership of reliable cloud-based systems.
Build and scale post-training, reinforcement-learning, evaluation, and inference systems for long-horizon agents operating over complex enterprise software. The role requires strong Python and PyTorch or JAX skills, distributed GPU experience, empirical rigor, and the ability to take research results into production.
Build and operate production AI agents that transform enterprise processes, data, and code. The role focuses on tool layers, retrieval, context management, evaluations, monitoring, auditability, and guardrails, requiring strong Python and TypeScript plus experience with production LLM systems and traditional machine learning.
Build and productionize applied AI/ML systems for document understanding, agentic workflows, and demand forecasting using rich, messy enterprise data. The role requires 3+ years of production AI/ML experience, strong evaluation and monitoring practices, and a STEM master’s degree.