Copy of Machine Learning Researcher, Audio
Conducts foundational research and develops scalable ML models for speech-to-text, text-to-speech, and neural audio codecs in real-time voice AI agents. Requires deep expertise in voice modeling, self-supervised learning, and production deployment at enterprise scale.
About the job
What You Will Do
Build and Scale Next-Generation TTS Systems
- Design and train large scale text-to-speech models capable of expressive, controllable, human-sounding output.
- Develop neural audio codec-based TTS architectures for efficient, high-fidelity generation.
- Improve prosody modeling, question inflection, emotional expression, and multi-speaker robustness.
- Optimize for real-time, low-latency inference in production.
Advance Speech-to-Text Modeling
- Build and fine-tune large scale ASR systems robust to accents, noise, telephony artifacts, and code switching.
- Leverage self-supervised pretraining and large-scale weak supervision.
- Improve transcription accuracy for real-world enterprise scenarios, including structured extraction and conversational nuance.
Pioneer Neural Audio Codecs
- Research and implement neural audio codecs that achieve extreme compression with minimal perceptual loss.
- Explore discrete and continuous latent representations for scalable speech modeling.
- Design codec architectures that enable downstream generative modeling and controllable synthesis.
Develop Scalable Training Pipelines
- Curate and process massive audio datasets across languages, speakers, and environments.
- Design staged training curricula and data filtering strategies.
- Scale training across distributed GPU clusters focusing on cost, throughput, and reliability.
Run Rigorous Experiments
- Design ablation studies that isolate the impact of architectural changes.
- Measure improvements using both objective metrics and perceptual evaluations.
- Validate ideas quickly through focused experiments that confirm or eliminate hypotheses.
What Makes You a Great Fit
Deep Research Foundations
- Experience with self-supervised learning, multimodal modeling, or generative modeling.
- Ability to derive new formulations and implement them efficiently.
Expertise in Voice Modeling
- Hands-on experience building or scaling TTS, STT, or neural audio codec systems.
- Familiarity with large scale speech datasets and real-world audio variability.
- Strong intuition for audio quality, prosody, and conversational dynamics.
Systems and Hardware Awareness
- Experience training and serving large models on modern accelerators.
- Knowledge of inference optimization techniques, including quantization, kernel optimization, and memory efficiency.
- Understanding of real-time constraints in telephony or streaming environments.
Experimental Rigor
- Track record of designing controlled experiments and meaningful ablations.
- Comfortable working with both offline benchmarks and live production metrics.
- Ability to move quickly from hypothesis to validation.
Bonus Points
- Experience with large scale distributed training.
- Research publications or open source contributions in speech or language AI.
- Background in real-time speech systems or telephony.
- PhD in ML, AI, or a related field, or equivalent research impact.
Benefits and Compensation
- Healthcare, dental, vision.
- Meaningful equity in a fast-growing company.
- Every tool you need to succeed.
- Beautiful office in Jackson Square, SF with rooftop views.
- Competitive salary: $160,000 to $250,000.
Skills
Text-To-Speech, Tts, Speech-To-Text, Stt, Neural Audio Codecs, Self-Supervised Learning, Generative Modeling, Asr, PyTorch, Distributed Training, Gpu Training, Inference Optimization
Similar jobs
AI Research jobsResearch Engineer focused on designing benchmarks, evaluation systems, rubrics, and failure-analysis workflows for frontier language models. The role requires strong applied AI research and coding experience, with expertise in model evaluation, data quality, and backend systems.
Conduct research on long-horizon, multi-agent AI behavior by designing agent environments, analyzing large-scale data, and running experiments. The role requires strong research judgment, rapid execution, independence, and familiarity with current AI developments.
Build, optimize, and evaluate long-running and multi-agent AI systems, along with tools for monitoring and analyzing their real-world behavior. The role requires software engineering experience with coding agents, strong independence, and familiarity with current AI developments.
Research Scientist developing and evaluating health-focused AI models, large language models, and agentic systems for clinical applications. The role requires advanced research experience, strong coding skills, healthcare or clinical-data experience, and top-tier AI/ML publications.
Develops experimental AI techniques and prototypes for agentic marketing applications, with emphasis on image and video generation. The role requires strong backend or probabilistic systems expertise, quantitative thinking, creativity with LLM applications, and product intuition.