Audio Inference Engineer, Model Efficiency
Develops high-performance audio inference systems, optimizing latency, throughput, and quality for real-time streaming workloads. Requires expertise in C++, Python, and deep learning models for audio/speech, with collaboration across training and serving teams.
About the job
Responsibilities
- Advance core audio model serving metrics, including latency, throughput, and quality.
- Dive deep into systems, identify bottlenecks, and deliver creative solutions for audio processing and streaming workloads.
- Collaborate closely with training and serving infrastructure teams for seamless integration between model development and deployment.
- Special focus on real-time and streaming audio inference.
Requirements
- Significant experience developing high-performance audio or machine learning inference systems.
- Proficiency with programming languages such as C++ and Python.
- Hands-on experience with deep learning models for audio, speech, or language applications.
- Bias for action and strong results-oriented mindset.
Nice-to-Haves
- GPU programming, low-level system optimization, model parallelization techniques over multiple GPUs.
- Experience with duplex real-time streaming architectures.
- Internals of machine learning frameworks for audio (PyTorch, TensorFlow, or specialized audio libraries).
- Experience with inference frameworks like vLLM, SGLang, TensorRT-LLM, or custom distributed inference systems.
- Sequence modeling (e.g., transformers for audio/speech) and end-to-end audio pipeline optimization.
Skills
C++, Python, PyTorch, TensorFlow, Gpu Programming, vLLM, Sglang, Tensorrt-Llm, Deep Learning, Transformers
Similar jobs
ML Engineering jobsBuild and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.
Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.
Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.
Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.