# Research Engineer, Audio and Speech

**Company:** [Decagon](https://hotfix.jobs/companies/decagon)
**Location:** San Francisco, CA, New York, NY
**Role:** ML Engineering
**Salary:** $200k – $400k/yr
**Experience:** 2+ years
**Skills:** Python, PyTorch, Speech Recognition, Voice Activity Detection, Speech Generation, Multimodal Machine Learning, Deep Learning, Signal Processing, Low-Latency Inference, Model Serving, Autoregressive Models, Diffusion Models, Flow Matching, Full-Duplex Models, Telephony
**Posted:** 2026-09-04

> Research Engineer focused on building and deploying real-time audio and speech models for conversational voice agents. The role requires experience with speech or multimodal machine learning, production inference, Python, and PyTorch, with emphasis on taking research from prototype to measurable production impact.

## Job Description

## Responsibilities
- Design and build agent harnesses optimized for streaming speech, turn-taking, interruptions, overlapping speech, and continuous interaction.
- Research and train multimodal and full-duplex models that jointly understand audio, reason, and generate speech.
- Improve speech recognition, voice activity detection, endpointing, and speech generation across diverse speakers, environments, domains, and languages.
- Build evaluations and use production calls to deliver measurable improvements in accuracy, latency, naturalness, and task outcomes.
- Optimize end-to-end inference for responsiveness, throughput, stability, and cost, partnering with platform and infrastructure teams for deployment at scale.

## Requirements
- 2+ years of experience in speech, audio machine learning, multimodal machine learning, or production machine learning.
- Experience developing or adapting autoregressive, diffusion, flow-matching, or codec-based speech models.
- Hands-on experience with streaming agent systems, low-latency inference, production model serving, and evaluation on real-world audio.
- Fluency in Python and a modern deep-learning framework such as PyTorch, with strong foundations in machine learning and signal processing.
- Track record of taking research ideas from prototype to reliable, measurable production impact.

## Nice-to-haves
- Familiarity with speech-to-speech or full-duplex models.
- Experience with telephony, multilingual speech, noisy-channel robustness, speaker adaptation, or expressive speech generation.

## Compensation and Benefits
- $200,000–$400,000 plus equity.
- Medical, dental, and vision benefits for employees and families.
- Life insurance and disability benefits.
- Retirement plan.
- Parental leave.
- Fertility and family-building benefits.
- Monthly wellness and lifestyle stipend.
- Daily office lunches and snacks.
- Flexible vacation policy.

## Similar jobs

- [Research Engineer, Safety](https://hotfix.jobs/jobs/6c559226-e2b7-4fd9-8c55-c26cb8e3dfe2) - Decagon - San Francisco, CA - $200k – $400k/yr
- [Research Engineer](https://hotfix.jobs/jobs/8d2c993e-0579-4dcc-bd83-6020a652dde9) - AfterQuery - San Francisco, CA - $210k – $450k/yr
- [Machine Learning Engineer](https://hotfix.jobs/jobs/9dfd52c1-58d4-4136-b330-ca02613097f2) - Stripe - South San Francisco, CA - $212k – $318k/yr
- [Machine Learning Engineer](https://hotfix.jobs/jobs/d575376e-3c50-47c4-8f40-6f254583ab42) - Earnin - Mountain View, CA - $187k – $229k/yr
- [Applied ML Scientist, New Grad](https://hotfix.jobs/jobs/adfb42c1-5eb0-4d5a-bd6a-582c98bca4a9) - SentiLink - Remote - $180k – $220k/yr

**Apply:** https://hotfix.jobs/jobs/728e9b38-e55c-473f-a4b8-7b26bb028516
**Canonical:** https://hotfix.jobs/jobs/728e9b38-e55c-473f-a4b8-7b26bb028516