Staff Machine Learning Engineer, Voice AI
Staff ML Engineer to own the model serving stack for real-time voice inference (STT, TTS, speech-to-speech) on H100/H200 GPUs. Drive latency/throughput optimization using TRT-LLM and SGLang for models like Whisper and Parakeet.
About the job
Responsibilities
- Own the voice inference roadmap end-to-end — define and execute the technical strategy for optimizing STT, TTS, and speech-to-speech models across Together's infrastructure
- Drive best-in-class inference performance — architect and implement systems targeting leading TTFB, throughput, and GPU utilization for voice workloads
- Lead productionization of voice models at scale — design the serving architecture for serverless and dedicated endpoints, including batching strategies, streaming inference pipelines, and memory management tailored to real-time audio
- Build the voice evaluation platform — design a rigorous evaluation framework covering WER across accents, languages, and noise conditions for STT; naturalness, latency, and pronunciation fidelity for TTS
- Shape the architecture for next-generation model support — anticipate and enable emerging model paradigms (audio-native LLMs, codec-based architectures, end-to-end speech-to-speech systems)
- Serve as the technical DRI for model partner integrations — lead collaboration with partners such as Cartesia, Deepgram, and Rime
- Diagnose and resolve performance problems — conduct systematic profiling and root-cause analysis from GPU kernel behavior to framework-level bottlenecks
- Influence platform architecture — partner with platform engineering leadership to ensure the serving layer meets latency and reliability demands of real-time voice APIs
- Define and scale voice fine-tuning capabilities — lead technical direction for enabling customers to fine-tune STT and TTS models on Together's infrastructure
Requirements
- 8+ years of ML engineering experience with focus on model serving, inference optimization, or ML infrastructure at production scale
- Deep expertise in LLM serving engines (vLLM, SGLang, TensorRT-LLM)
- Expert-level Python and PyTorch proficiency with strong command of GPU optimization (CUDA kernels, memory hierarchies, profiling toolchains)
- Proven system design judgment and architectural decisions that held up at scale
- Strong technical leadership with high autonomy
- Sharp product intuition for developer tooling
- Strong foundation in speech and audio ML (ASR/TTS architectures, audio signal processing) preferred
- Familiarity with audio codec and tokenization schemes (SNAC, Encodec, DAC) is a plus
- Experience training or fine-tuning speech models at scale is an advantage
- Bachelor's or Master's in Computer Science, Electrical Engineering, or related field
Nice-to-Haves
- Experience modifying engine internals and contributing improvements back to serving frameworks
- Proven ability to move fast in ambiguous, early-stage environments
Skills
Python, PyTorch, Tensorrt-Llm, Sglang, vLLM, CUDA, Gpu Optimization, Asr, Tts, Speech-To-Text, Text-To-Speech, Audio Signal Processing, Snac, Encodec, Model Serving
Similar jobs
ML Engineering jobsLeads architecture and technical direction for agentic search systems combining LLMs, retrieval, and content-understanding pipelines for contract intelligence. The role requires 10+ years building production systems, deep search or LLM expertise, and strong cross-team technical leadership.
Leads development and integration of advanced maritime autonomy for USVs, UUVs, and cooperating UAVs, including motion planning, localization, safety, and multi-agent coordination. Requires staff-level technical leadership, substantial robotics experience, C++ and Python proficiency, and eligibility for a SECRET clearance.
Leads the engineering discipline for evaluating, testing, and monitoring production AI agents, while building scalable eval infrastructure and developer tooling. Requires 8+ years of production software experience, strong backend skills, and expertise with LLM evaluation and agentic systems.
Architects and operates production machine-learning systems that classify web and API traffic, detect bots and scrapers, and support real-time mitigation at internet edge latency. The role requires 9+ years of applied ML experience in adversarial domains and strong expertise in evaluation, data pipelines, and large-scale systems.
Leads the design and operation of reliable, scalable model infrastructure powering AI inference across multiple providers. Requires 7+ years of distributed-systems engineering experience, strong programming skills, and expertise in production reliability and cloud infrastructure.