Applied ML Engineer
Own the research-to-production pipeline at Deepgram, turning experimental speech ML models into reliable, scalable production services. Partner with researchers on robust workflows, automated release gates, inference optimization, and feedback loops across hybrid GPU infrastructure.
About the job
What You'll Do
- Own the research-to-production pipeline: take research checkpoints and turn them into production models, defining the repeatable path from a working result to a deployed, monitored, scaled service.
- Partner directly with research scientists to productionize new models — translating experimental training and evaluation code into robust, reproducible, well-tested workflows.
- Build and extend the tooling and abstractions that let researchers and engineers move models through training, evaluation, packaging, and deployment with minimal friction and maximal reproducibility.
- Design and own model release gates — automated evaluation, regression detection, and quality/latency/throughput checks that decide whether a model is ready to ship.
- Optimize models and serving for production: efficient inference, batching, memory and latency tuning, and the profiling work that turns a research model into something that performs economically at scale.
- Strengthen the build and delivery layer for models on our custom infrastructure, spanning our GPU compute and cloud environments, so that shipping a model is fast, safe, and observable.
- Establish benchmarking and validation that runs consistently from model development all the way through production, so performance and quality regressions are caught early.
- Build the feedback loop: instrument production model behavior, surface what's working and what isn't, and feed it back to research to accelerate the next iteration.
Requirements
- Strong software engineering fundamentals, with proficiency in Python and experience writing production-quality, well-tested ML code.
- Hands-on experience taking ML models from research or prototype stage into production at scale — not just training models, but shipping and operating them.
- A working understanding of the modern deep learning stack (e.g., PyTorch) and the realities of training, evaluating, and serving large models.
- Experience building ML pipelines and tooling — training orchestration, evaluation harnesses, model packaging, deployment, or CI/CD for models.
- Familiarity with serving and inference optimization — latency, throughput, batching, and resource efficiency for production model workloads.
- Comfort operating across distributed systems and GPU compute, whether in the cloud, on bare metal, or both.
- A collaborative, builder mindset — you can partner with researchers, scope an ambiguous problem, and drive it to a measurable result.
Nice-to-Haves
- Experience with the research-to-production handoff specifically — building the systems and conventions that let research and engineering iterate together quickly.
- Background in speech, audio, or other real-time/streaming ML domains.
- Experience designing automated model evaluation and release-gating systems, including regression detection across model versions.
- Familiarity with hybrid infrastructure spanning on-premise GPU clusters and cloud, and with workload orchestration across them.
- Experience with inference optimization techniques (quantization, distillation, compilation, or runtime tuning) for production serving.
- A track record of building internal platforms or developer-facing tooling that measurably improved how a team ships models.
Skills
Python, PyTorch, Ml Pipelines, Model Deployment, Inference Optimization, Gpu Compute, Distributed Systems, Ci/Cd For Ml, Model Evaluation, Quantization
Similar jobs
ML Engineering jobsBuild and teach reliable AI agent systems through customer workshops, technical content, guidance, and reference implementations. The role requires strong Python and agent-development experience plus a background delivering customer-facing technical training.
Build reinforcement-learning environments, evaluations, datasets, and scalable infrastructure for frontier AI capabilities. The role suits a high-agency generalist engineer with experience in agents, evaluations, or RL workflows and strong communication skills.
Develop and productionize machine- and deep-learning algorithms for biosignal and EEG data used in medical devices, clinical development, and diagnostics. The role requires 4+ years of industry experience, DSP and statistics expertise, PyTorch proficiency, and familiarity with regulated environments and production ML practices.
Develop and deploy ML-first behavior prediction and planning systems for autonomous vehicles, forecasting the motion and interactions of road users. Requires a bachelor's degree, deep learning lifecycle expertise, and at least three years of production software experience with C++ or Python.
Build the AI platform behind fab2, including model infrastructure, agent systems, evaluations, and tools for engineering and fab operations. The role requires strong production software engineering skills and comfort working across frontend, backend, infrastructure, and data.