Director of Research, Text to Speech
Leads Deepgram’s end-to-end TTS research program, setting technical direction, training and evaluating large-scale speech-generation models, and turning breakthroughs into production systems. The role combines hands-on technical leadership with building and developing a high-performing research organization.
About the job
Responsibilities
- Own the text-to-speech research and model roadmap, including technical direction, experiments, training strategy, and production goals.
- Advance neural audio modeling, prosody and expressiveness, controllability, multilingual speech, voice identity and consistency, data and training strategy, post-training, and inference performance.
- Review research, challenge assumptions, design experiments, diagnose model failures, and address high-leverage technical problems.
- Build evaluation and benchmarking systems combining automated metrics with human perceptual assessment.
- Lead individual contributors and tech lead managers; hire, develop, and set direction across research sub-teams.
- Partner with engineering and product leadership on ship-readiness and represent TTS research internally and externally.
Requirements
- Deep expertise in modern TTS, speech generation, or audio generative modeling.
- Track record of personally training and improving large-scale neural models.
- Strong command of speech-generation problems including naturalness, expressiveness, controllability, robustness, voice consistency, and inference cost.
- Experience setting research direction under uncertainty and prioritizing experiments, compute, and researcher time.
- Experience leading researchers and research engineers through technical leaders while remaining technically influential.
- AI-first working style and demonstrated ability to integrate AI into research workflows.
- Ability to explain complex technical tradeoffs to product, engineering, and executive audiences.
Nice-to-haves
- TTS or generative-audio models deployed at meaningful production scale.
- Experience building or scaling a high-performing AI research organization.
- Sophisticated evaluation systems for generative speech.
- Experience with expressive or multilingual generation, voice cloning and adaptation, or controllable generation.
- External contributions such as publications, open source, patents, or invited talks in speech synthesis, neural audio codecs, speech language models, or multimodal models.
- Experience taking models from research to production in startup or research environments.
Skills
Text-To-Speech, Speech Generation, Audio Generative Modeling, Neural Audio Modeling, Prosody, Multilingual Speech, Voice Cloning, Model Evaluation, Human Perceptual Assessment, Neural Audio Codecs, Speech Language Models, Multimodal Models, Deep Learning, Model Training, Inference Optimization
Similar jobs
AI Research jobsLeads the research agenda and hands-on development of replayable enterprise environments, agent evaluations, and post-training systems. The role requires deep AI research experience, a PhD or equivalent track record, and the ability to translate open-ended questions into production systems.
Sets company-wide architecture and strategy for data and applied AI, connecting governed data foundations to production intelligence and measurable business outcomes. The role requires 14+ years of experience, strong production engineering judgment, executive partnership, and hands-on delivery.
Leads technical direction and develops maritime autonomy capabilities for unmanned surface and underwater vehicles, including motion planning, localization, safe behaviors, and heterogeneous multi-agent collaboration. Requires deep robotics and unmanned-systems experience, strong C++/Python skills, and senior technical leadership.
Leads design, implementation, integration, and field validation of tactical autonomy software for unmanned systems and multi-agent missions. Requires extensive autonomy or robotics experience, strong C++ and Python skills, and the ability to obtain a SECRET clearance.
Evaluates model and Generative AI risks across Upstart Bank’s model inventory, conducting risk assessments, monitoring reviews, quantitative analyses, and governance activities. Requires a quantitative master’s degree, 4+ years of relevant experience, and coding skills in Python, R, or similar languages.