Embedded AI Engineer, On-Device Models
Optimize and deploy Deepgram's state-of-the-art speech AI models onto resource-constrained embedded devices, edge hardware, and consumer products. Requires strong C/C++/Rust skills, on-device ML optimization experience, and deep hardware-software co-design knowledge for low-power, real-time inference.
About the job
What You'll Do
- Take Deepgram's Speech and Conversational models and get them running on embedded and low-power consumer hardware — defining the architecture for on-device, real-time inference across a diverse range of processors and accelerators.
- Optimize models for constrained targets through quantization, pruning, distillation, operator fusion, and architecture-specific compilation to meet strict latency, memory, power, and thermal budgets.
- Write and optimize performance-critical runtime code (C, C++, and/or Rust) for embedded environments, including bare-metal and real-time operating systems such as FreeRTOS and Zephyr.
- Integrate with industry-standard edge inference runtimes and vendor NPU/DSP toolchains, mapping model graphs efficiently onto on-device accelerators and CPU/GPU/NPU heterogeneity.
- Build the on-device runtime plumbing: model packaging, deployment pipelines, over-the-air update mechanisms, and lightweight telemetry for devices operating with limited or intermittent connectivity.
- Establish repeatable benchmarking and validation across target hardware — measuring latency, accuracy, power consumption, memory footprint, and resource utilization — and catch regressions before they ship.
- Partner with silicon and device vendors on SDK integration and performance tuning, getting our models to run efficiently on new chipsets and reference platforms.
- Collaborate with Research and Engine teams to influence model architectures toward edge-friendly designs from the start, reducing the optimization burden at deployment time.
Requirements
- Experience delivering production systems on resource-constrained hardware — embedded systems, mobile, edge AI, or small low-power devices.
- Strong proficiency in C, C++, and/or Rust, with experience writing performance-critical code for constrained environments.
- Hands-on experience with model optimization for on-device deployment, including quantization, pruning, knowledge distillation, or architecture-specific compilation.
- Familiarity with edge inference runtimes (e.g., ONNX Runtime, TensorRT, TFLite, ExecuTorch) and/or vendor-specific NPU/DSP toolchains.
- A strong understanding of hardware-software interaction — CPU/GPU/NPU/DSP architectures, memory hierarchies, fixed-point/integer arithmetic, and power management — and how they affect inference performance.
- Experience working close to the metal: bare-metal or RTOS environments (e.g., FreeRTOS, Zephyr), embedded Linux, or microcontroller and edge SoC development.
- Strong communication skills and a builder mindset — you can scope an ambiguous optimization problem, drive it to a measurable result, and explain the tradeoffs clearly.
Nice-to-Haves
- Experience with real-time audio processing on embedded platforms — DSP pipelines, audio codec optimization, wake-word or always-on listening, or streaming inference on microcontrollers and edge SoCs.
- Depth in ML optimization techniques — custom quantization schemes, mixed-precision inference, or neural architecture search for edge targets.
- Background in hardware evaluation and benchmarking — systematically comparing accelerators, SoCs, or GPUs for specific workload profiles.
- Experience shipping AI features in consumer products at scale, and the instinct for what "production quality" means on a battery-powered device.
- Familiarity with model compilation and optimization toolchains and their tradeoffs across hardware targets.
- Experience with secure, robust on-device deployment practices — code signing, encrypted model storage, and safe update mechanisms.
Skills
C++, Rust, Quantization, Pruning, Distillation, Onnx Runtime, TensorRT, Tflite, Executorch, Freertos, Zephyr, Npu, Dsp, Embedded Linux
Similar jobs
Embedded Engineering jobsBuild and harden OS foundations for AI consumer devices, spanning kernel, services, security, performance, and application interfaces. Requires strong systems programming in C/C++, kernel experience, and debugging skills.
Develops real-time perception, sensor-control, and fusion software for heterogeneous autonomous vehicle platforms, integrating ML algorithms across sensors and embedded systems. Requires advanced engineering education or 5+ years of relevant experience, multi-modal sensing expertise, Linux/Docker proficiency, and U.S. security-clearance eligibility.
Develops operating-system platforms, device drivers, and infrastructure for autonomous-vehicle robotics systems. The role requires 4+ years of production software experience, strong Linux internals expertise, and proficiency in C/C++ plus scripting languages.
Develops reliable embedded software and custom device solutions, integrating vendor components, sensors, operating systems, and communication interfaces. Requires a bachelor's degree and at least three years of relevant experience with C/C++, embedded Linux, debugging, and CI/CD tools.
Build the low-level runtime for a custom AI accelerator, including kernel scheduling, device memory management, synchronization, and hardware-software interfaces. The role requires strong systems programming experience and expertise in concurrency, memory semantics, simulation, and performance debugging.