AI Inference Engineer - Model Optimization & Deployment
Optimizes and deploys large-scale AI models (LLMs, VLMs) for real-time inference on power-constrained vehicle hardware. Requires expertise in quantization, TensorRT compilation, custom CUDA kernels, and production C++/Python for edge devices.
About the job
Responsibilities
- Optimize large-scale models (LLMs, VLMs) using advanced quantization (PTQ, QAT), mixed-precision inference workflows, and parameter-efficient fine-tuning (LoRA, QLoRA).
- Architect and implement model conversion and compilation pipelines using TensorRT and TensorRT-LLM for edge deployment.
- Perform rigorous parity checking, accuracy recovery, and latency benchmarking between PyTorch frameworks and compiled edge binaries.
- Write and optimize custom CUDA kernels and TensorRT Plugins to maximize memory bandwidth and minimize latency on AI accelerators.
- Write production-level, highly concurrent, and memory-safe C++ and Python code for real-time inference on vehicle SOCs.
Qualifications
- Deep expertise in model quantization (PTQ, QAT) and mixed-precision inference workflows (INT8, FP8, INT4, BF16/FP16).
- Proven experience optimizing large-scale models (LLMs, VLMs, or VLAs) utilizing KV-cache optimization (e.g., PagedAttention), Speculative Decoding, and Efficient Attention mechanisms (FlashAttention, Linear Attention).
- Extensive experience with model conversion/compilation pipelines (TensorRT, TensorRT-LLM) and performing rigorous parity/latency benchmarking.
- Proficiency in low-level programming for AI accelerators, specifically writing and optimizing custom CUDA kernels and TensorRT Plugins.
- Production-level C++ (14/17/20) and Python programming skills, with experience writing concurrent, memory-safe, real-time inference code for edge devices.
Bonus Qualifications
- Experience with distributed training pipelines and model/tensor parallelism (PyTorch Distributed, Ray, DeepSpeed, Megatron-LM) and runtime efficiency optimization for GPU clusters.
- Familiarity with autonomous driving perception stacks (temporal 3D object detection, BEV, 3D Occupancy Networks) and processing multi-modal sensor streams (Vision, LiDAR, Radar).
- Understanding of end-to-end autonomous driving paradigms (VLA models, closed-loop simulation validation).
Skills
TensorRT, Tensorrt-Llm, CUDA, PyTorch, C++, Python, Ptq, Qat, Lora, Qlora, Flashattention, Pagedattention, Deepspeed, Ray
Similar jobs
ML Engineering jobsBuild and deploy algorithmic systems for high-impact healthcare problems, choosing among machine learning, optimization, heuristics, and hybrid approaches. The role requires 4+ years of relevant industry experience, strong applied problem-solving and evaluation skills, and fluency in modern ML tooling.
Build and operate machine learning models for sales roleplay, scoring, and coaching products, owning the lifecycle from fine-tuning and evaluation through production and on-device deployment. The role emphasizes open-source models, latency and privacy optimization, and rigorous model testing.
Build and operate production machine-learning systems for content safety, from messy customer data through classification, evaluation, and inference. The role requires 5+ years of ML engineering experience, strong Python and MLOps skills, and sound judgment across classical models and LLMs.
Build AI agent harnesses, models, and product capabilities that enable agents to perform complex work across digital environments. The role combines applied AI research and software engineering, requiring Python proficiency, strong product judgment, and experience with agent tooling, reinforcement learning, or browser technologies.
Builds the platform, verifiers, environments, and grading infrastructure used to evaluate enterprise AI agents at scale. The role combines strong software engineering with expertise in agent runtimes, evaluation design, benchmarks, and production failure analysis.