Senior Machine Learning Engineer, Voice Agents
Own and evolve Hugging Face’s open-source voice-agent stack and bring hf-voice from demo to production. The role requires senior experience with Python, distributed real-time systems, developer infrastructure, and production AI or multimodal models.
About the job
Responsibilities
- Own the architecture of the open-source
speech-to-speechlibrary, including pipeline design, latency budgets, and real-time loop reliability. - Integrate new automatic speech recognition (ASR), text-to-speech (TTS), and end-to-end speech models while maintaining clean abstractions.
- Review community pull requests, triage issues, release versions, and grow the contributor community.
- Design the
hf-voicedeveloper API and streaming protocol, including session lifecycle, WebSockets/WebRTC transport, authentication, error semantics, and versioning. - Build real-time GPU inference serving, concurrency, autoscaling, observability, and cost controls.
- Collaborate with Hub and inference teams to integrate voice agents into products and demonstrations.
- Take the product from demo to production through load testing, SLOs, and graceful degradation.
- Write documentation, examples, and templates; support existing deployments such as the Reachy Mini fleet.
- Optionally present the work through blog posts, demos, and conference talks.
Requirements
- Senior-level experience owning substantial architecture and driving projects autonomously.
- Experience building developer-facing infrastructure at an AI or developer-tools company, such as inference APIs or agent infrastructure.
- Substantial open-source contributions to a Python library.
- Proficiency with asynchronous Python and distributed systems, including failure modes.
- Experience shipping real-time systems involving streaming, WebSockets, WebRTC, audio/video pipelines, or live inference.
- Production experience with LLMs or multimodal models.
- Clear written communication and experience collaborating asynchronously and in public.
- Motivation to work on voice and conversational AI.
Nice-to-haves
- Contributions to voice-agent frameworks such as speech-to-speech, Pipecat, LiveKit Agents, Vocode, or TEN.
- Contributions to
llama.cppor another low-level inference runtime. - Experience with ASR, TTS, or end-to-end speech models, including latency and quality evaluation.
- GPU serving, quantization, or on-device inference experience.
- Audio pipeline knowledge, including voice activity detection, echo cancellation, jitter buffers, barge-in, and turn detection.
- Experience shipping to embedded or robotics targets.
- Public technical work such as talks, blog posts, or demos.
Compensation and Benefits
- Company equity.
- Conference, training, and education reimbursement.
- Health, dental, and vision benefits for employees and dependents.
- Parental leave and flexible paid time off.
- Flexible working hours and remote work options.
- Distributed-work support, office access, and workstation equipment when needed.
Skills
Python, Async Python, Distributed Systems, WebSockets, Webrtc, Gpu Serving, LLMs, Multimodal Models, Automatic Speech Recognition, Text-To-Speech, Quantization, Voice Activity Detection, Audio Pipelines, Real-Time Inference, Autoscaling
Similar jobs
ML Engineering jobsBuild and operate AI-powered, customer-facing workflows for Datadog Notebooks, combining reliable backend systems with LLM capabilities. The role requires 6+ years of engineering experience, Go or Python expertise, and experience delivering production AI products.
Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.
Build the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.
Develop and productionize machine- and deep-learning algorithms for biosignal and EEG data used in medical devices, clinical development, and diagnostics. The role requires 4+ years of industry experience, DSP and statistics expertise, PyTorch proficiency, and familiarity with regulated environments and production ML practices.
Build production agent systems that plan, use tools, recover from failures, and improve over time. The role requires 5+ years of production ML or backend experience, LLM or agent deployment experience, and expertise in evaluation, tracing, observability, and agent architecture.