Skip to content
DeepgramDeepgram

Embedded AI Engineer, On-Device Models

Optimize and deploy Deepgram's state-of-the-art speech AI models onto resource-constrained embedded devices, edge hardware, and consumer products. Requires strong C/C++/Rust skills, on-device ML optimization experience, and deep hardware-software co-design knowledge for low-power, real-time inference.

About the job

What You'll Do

  • Take Deepgram's Speech and Conversational models and get them running on embedded and low-power consumer hardware — defining the architecture for on-device, real-time inference across a diverse range of processors and accelerators.
  • Optimize models for constrained targets through quantization, pruning, distillation, operator fusion, and architecture-specific compilation to meet strict latency, memory, power, and thermal budgets.
  • Write and optimize performance-critical runtime code (C, C++, and/or Rust) for embedded environments, including bare-metal and real-time operating systems such as FreeRTOS and Zephyr.
  • Integrate with industry-standard edge inference runtimes and vendor NPU/DSP toolchains, mapping model graphs efficiently onto on-device accelerators and CPU/GPU/NPU heterogeneity.
  • Build the on-device runtime plumbing: model packaging, deployment pipelines, over-the-air update mechanisms, and lightweight telemetry for devices operating with limited or intermittent connectivity.
  • Establish repeatable benchmarking and validation across target hardware — measuring latency, accuracy, power consumption, memory footprint, and resource utilization — and catch regressions before they ship.
  • Partner with silicon and device vendors on SDK integration and performance tuning, getting our models to run efficiently on new chipsets and reference platforms.
  • Collaborate with Research and Engine teams to influence model architectures toward edge-friendly designs from the start, reducing the optimization burden at deployment time.

Requirements

  • Experience delivering production systems on resource-constrained hardware — embedded systems, mobile, edge AI, or small low-power devices.
  • Strong proficiency in C, C++, and/or Rust, with experience writing performance-critical code for constrained environments.
  • Hands-on experience with model optimization for on-device deployment, including quantization, pruning, knowledge distillation, or architecture-specific compilation.
  • Familiarity with edge inference runtimes (e.g., ONNX Runtime, TensorRT, TFLite, ExecuTorch) and/or vendor-specific NPU/DSP toolchains.
  • A strong understanding of hardware-software interaction — CPU/GPU/NPU/DSP architectures, memory hierarchies, fixed-point/integer arithmetic, and power management — and how they affect inference performance.
  • Experience working close to the metal: bare-metal or RTOS environments (e.g., FreeRTOS, Zephyr), embedded Linux, or microcontroller and edge SoC development.
  • Strong communication skills and a builder mindset — you can scope an ambiguous optimization problem, drive it to a measurable result, and explain the tradeoffs clearly.

Nice-to-Haves

  • Experience with real-time audio processing on embedded platforms — DSP pipelines, audio codec optimization, wake-word or always-on listening, or streaming inference on microcontrollers and edge SoCs.
  • Depth in ML optimization techniques — custom quantization schemes, mixed-precision inference, or neural architecture search for edge targets.
  • Background in hardware evaluation and benchmarking — systematically comparing accelerators, SoCs, or GPUs for specific workload profiles.
  • Experience shipping AI features in consumer products at scale, and the instinct for what "production quality" means on a battery-powered device.
  • Familiarity with model compilation and optimization toolchains and their tradeoffs across hardware targets.
  • Experience with secure, robust on-device deployment practices — code signing, encrypted model storage, and safe update mechanisms.

Skills

C++, Rust, Quantization, Pruning, Distillation, Onnx Runtime, TensorRT, Tflite, Executorch, Freertos, Zephyr, Npu, Dsp, Embedded Linux

OpenAI

OpenAI

San Francisco, CA

Operating Systems Engineer | Consumer Devices
$230k+/yrOn-siteEmbedded Engineering

Build and harden OS foundations for AI consumer devices, spanning kernel, services, security, performance, and application interfaces. Requires strong systems programming in C/C++, kernel experience, and debugging skills.

Applied Intuition

Applied Intuition

Sunnyvale, CA

Perception & Fusion Engineer
$200k+/yrOn-site5+ YOEEmbedded Engineering

Develops real-time perception, sensor-control, and fusion software for heterogeneous autonomous vehicle platforms, integrating ML algorithms across sensors and embedded systems. Requires advanced engineering education or 5+ years of relevant experience, multi-modal sensing expertise, Linux/Docker proficiency, and U.S. security-clearance eligibility.

Zoox

Zoox

Foster City, CA

Software Engineer - Robot Software Infrastructure
$196k+/yrHybrid4+ YOEEmbedded Engineering

Develops operating-system platforms, device drivers, and infrastructure for autonomous-vehicle robotics systems. The role requires 4+ years of production software experience, strong Linux internals expertise, and proficiency in C/C++ plus scripting languages.

Applied Intuition

Applied Intuition

Sunnyvale, CA

Software Engineer
$188k+/yrOn-site3+ YOEEmbedded Engineering

Develops reliable embedded software and custom device solutions, integrating vendor components, sensors, operating systems, and communication interfaces. Requires a bachelor's degree and at least three years of relevant experience with C/C++, embedded Linux, debugging, and CI/CD tools.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, AI Accelerator Runtime
$266k+/yrHybridEmbedded Engineering

Build the low-level runtime for a custom AI accelerator, including kernel scheduling, device memory management, synchronization, and hardware-software interfaces. The role requires strong systems programming experience and expertise in concurrency, memory semantics, simulation, and performance debugging.