Skip to content
ZooxZoox

AI Inference Engineer - Model Optimization & Deployment

Optimizes and deploys large-scale AI models (LLMs, VLMs) for real-time inference on power-constrained vehicle hardware. Requires expertise in quantization, TensorRT compilation, custom CUDA kernels, and production C++/Python for edge devices.

About the job

Responsibilities

  • Optimize large-scale models (LLMs, VLMs) using advanced quantization (PTQ, QAT), mixed-precision inference workflows, and parameter-efficient fine-tuning (LoRA, QLoRA).
  • Architect and implement model conversion and compilation pipelines using TensorRT and TensorRT-LLM for edge deployment.
  • Perform rigorous parity checking, accuracy recovery, and latency benchmarking between PyTorch frameworks and compiled edge binaries.
  • Write and optimize custom CUDA kernels and TensorRT Plugins to maximize memory bandwidth and minimize latency on AI accelerators.
  • Write production-level, highly concurrent, and memory-safe C++ and Python code for real-time inference on vehicle SOCs.

Qualifications

  • Deep expertise in model quantization (PTQ, QAT) and mixed-precision inference workflows (INT8, FP8, INT4, BF16/FP16).
  • Proven experience optimizing large-scale models (LLMs, VLMs, or VLAs) utilizing KV-cache optimization (e.g., PagedAttention), Speculative Decoding, and Efficient Attention mechanisms (FlashAttention, Linear Attention).
  • Extensive experience with model conversion/compilation pipelines (TensorRT, TensorRT-LLM) and performing rigorous parity/latency benchmarking.
  • Proficiency in low-level programming for AI accelerators, specifically writing and optimizing custom CUDA kernels and TensorRT Plugins.
  • Production-level C++ (14/17/20) and Python programming skills, with experience writing concurrent, memory-safe, real-time inference code for edge devices.

Bonus Qualifications

  • Experience with distributed training pipelines and model/tensor parallelism (PyTorch Distributed, Ray, DeepSpeed, Megatron-LM) and runtime efficiency optimization for GPU clusters.
  • Familiarity with autonomous driving perception stacks (temporal 3D object detection, BEV, 3D Occupancy Networks) and processing multi-modal sensor streams (Vision, LiDAR, Radar).
  • Understanding of end-to-end autonomous driving paradigms (VLA models, closed-loop simulation validation).

Skills

TensorRT, Tensorrt-Llm, CUDA, PyTorch, C++, Python, Ptq, Qat, Lora, Qlora, Flashattention, Pagedattention, Deepspeed, Ray

Garner Health

Garner Health

New York, NY

Applied Scientist III
$236k+/yrOn-site4+ YOEML Engineering

Build and deploy algorithmic systems for high-impact healthcare problems, choosing among machine learning, optimization, heuristics, and hybrid approaches. The role requires 4+ years of relevant industry experience, strong applied problem-solving and evaluation skills, and fluency in modern ML tooling.

Hyperbound

Hyperbound

San Francisco, CA

Machine Learning Engineer
$260k+/yrOn-siteML Engineering

Build and operate machine learning models for sales roleplay, scoring, and coaching products, owning the lifecycle from fine-tuning and evaluation through production and on-device deployment. The role emphasizes open-source models, latency and privacy optimization, and rigorous model testing.

Cinder

Cinder

New York, NY

AI/ML Engineer
$220k+/yrHybrid5+ YOEML Engineering

Build and operate production machine-learning systems for content safety, from messy customer data through classification, evaluation, and inference. The role requires 5+ years of ML engineering experience, strong Python and MLOps skills, and sound judgment across classical models and LLMs.

Perplexity

Perplexity

San Francisco, CA

Member of Technical Staff
$220k+/yrOn-siteML Engineering

Build AI agent harnesses, models, and product capabilities that enable agents to perform complex work across digital environments. The role combines applied AI research and software engineering, requiring Python proficiency, strong product judgment, and experience with agent tooling, reinforcement learning, or browser technologies.

Mercor

Mercor

San Francisco, CA

Member of Technical Staff, Enterprise Evals Platform
$220k+/yrOn-siteML Engineering

Builds the platform, verifiers, environments, and grading infrastructure used to evaluate enterprise AI agents at scale. The role combines strong software engineering with expertise in agent runtimes, evaluation design, benchmarks, and production failure analysis.