Senior AI Engineer
Leads development of speech models, decoders, and low-latency inference systems for next-generation voice agents. Requires 5+ years in speech ML or related audio AI, strong Python and PyTorch experience, and the ability to guide technical direction and mentor engineers.
About the job
Responsibilities
- Set technical direction for speech-model and decoder strategy across the team.
- Lead research, adaptation, and implementation of ASR/STT, speech-enhancement, and audio models.
- Improve end-to-end real-time speech pipelines, including decoding, endpointing, turn detection, streaming behavior, and ASR/LLM/TTS handoffs.
- Benchmark, prototype, fine-tune, distill, and adapt models and algorithms to improve voice-agent quality, latency, cost, and reliability.
- Design and ship production-grade services and inference components for real-time speech.
- Partner with speech, NLP, telephony, platform, product, and infrastructure engineers.
- Lead technical reviews, mentor teammates, and translate speech advances into measurable product improvements.
Requirements
- 5+ years of experience in speech ML, speech recognition, speech enhancement, audio AI, or a closely related field.
- Strong Python programming skills and experience with deep learning frameworks such as PyTorch.
- Hands-on experience improving speech models or systems used in real-world applications.
- Deep understanding of modern ASR/STT architectures and decoding techniques.
- Experience evaluating and adapting models across accuracy, robustness, streaming behavior, latency, and compute cost.
- Track record of turning research ideas, papers, experiments, or emerging models into measurable product improvements.
- Experience building, deploying, and operating low-latency, production-grade ML or backend services in a cloud environment.
- Ability to set technical direction, communicate across disciplines, mentor engineers, and balance model quality, reliability, latency, and customer impact.
Nice to Have
- Experience with streaming inference and observability.
- Experience with Google Cloud Platform (GCP).
- Background in research-heavy ML, backend/inference engineering, or both.
Compensation
- California target base salary: $224,500–$256,000 USD.
Skills
Python, PyTorch, Asr, Speech Recognition, Speech Enhancement, Audio Ai, Speech Models, Decoding, Real-Time Inference, Streaming Inference, Deep Learning, GCP, Backend Services, Observability
Similar jobs
ML Engineering jobsBuild and improve production AI systems for clinical products, owning evaluations, model behavior, agentic workflows, data flywheels, deployment, and observability. The role requires 5+ years of production ML or applied AI experience, strong Python and modern ML framework skills, and hands-on debugging expertise.
Senior AI Engineer responsible for production LLM agents that enrich business identity data through web discovery, verification, classification, and risk scoring. The role requires strong asynchronous Python, agent and evaluation expertise, browser automation, and experience operating AI systems in production.
Leads the development and production deployment of large-scale ASR and TTS systems for conversational intelligence products. The role requires 5+ years of industry experience, deep speech-model expertise, and strong software engineering and ML operations capabilities.
Design, build, and deploy production ML systems for recommendations, search, ranking, and advertising at internet scale. Own the full ML lifecycle from modeling to monitoring with strong cross-functional collaboration.
Build and operate low-latency machine learning systems for ad ranking, relevance, and optimization, including feature pipelines, experimentation, evaluation, and production inference. The role requires 6+ years of software engineering experience, strong Python skills, AWS experience, and practical LLM application experience.