Senior/Staff Software Engineer
Optimize and deploy large multi-modal foundation models (LLMs, VLMs) for real-time inference on power-constrained vehicle SoCs. Focus on quantization, custom CUDA kernels, TensorRT pipelines, resource allocation, and low-latency concurrent C++ code for autonomous systems.
About the job
Responsibilities
- Allocate and distribute system resources (CPU/GPU/interconnect) to various models and inference engines running on the robot.
- Spearhead cross-cutting initiatives that allow for better compute utilization through sharing/fusing models and better scheduling strategies.
- Optimize large-scale models (Multi-Modal Sensor Fusion models, LLMs, VLMs) using advanced quantization (PTQ, QAT), pruning, mixed-precision inference frameworks, and parameter-efficient fine-tuning (LoRA, QLoRA).
- Architect and implement model conversion and compilation pipelines using TensorRT for edge deployment.
- Write production-level, low-latency, and memory-safe C++ and CUDA code for real-time inference on vehicle systems.
Requirements
- Deep experience in system and performance optimization in CPU/GPU systems designed for low latency or high throughput.
- Deep expertise in working with real-time systems & required constraints such as processing latency, memory utilization, and memory bandwidth pressure.
- Deep expertise in model quantization (PTQ, QAT) and mixed-precision inference frameworks (INT8, FP8, FP4, BF16/FP16).
- Proficiency in low-level programming for AI accelerators, specifically developing and optimizing custom ML OPs and TensorRT Plugins with efficient CUDA kernel implementations.
- Production-level C++ (14/17/20) and Python programming skills, with experience developing concurrent, memory-safe, real-time inference code for edge devices.
Nice-to-Haves
- Prior experience in high-performance robotics applications such as AV/drones/robots.
- Familiarity with SOTA autonomous driving perception algorithms (temporal 3D object detection, BEV, 3D Occupancy Networks) and multi-modal sensor processing (Vision, LiDAR, Radar).
- Experience with end-to-end autonomous driving paradigms (VLM/VLA models, Foundation models) and edge deployment technologies (e.g., TensorRT-LLM).
Skills
CUDA, TensorRT, C++, Python, Model Quantization, Ptq, Qat, Mixed-Precision Inference, Lora, Qlora, Gpu Optimization, Real-Time Systems, Llm Optimization, Vlm Optimization
Similar jobs
ML Engineering jobsLeads the design and operation of reliable, scalable model infrastructure powering AI inference across multiple providers. Requires 7+ years of distributed-systems engineering experience, strong programming skills, and expertise in production reliability and cloud infrastructure.
Leads development and integration of advanced maritime autonomy for USVs, UUVs, and cooperating UAVs, including motion planning, localization, safety, and multi-agent coordination. Requires staff-level technical leadership, substantial robotics experience, C++ and Python proficiency, and eligibility for a SECRET clearance.
Leads architecture and technical direction for agentic search systems combining LLMs, retrieval, and content-understanding pipelines for contract intelligence. The role requires 10+ years building production systems, deep search or LLM expertise, and strong cross-team technical leadership.
Leads the design, implementation, integration, and field validation of tactical autonomy and multi-agent coordination capabilities for unmanned platforms. Requires 7+ years of relevant experience, production C++, technical leadership, and eligibility for a U.S. Secret clearance.
Leads the engineering discipline for evaluating, testing, and monitoring production AI agents, while building scalable eval infrastructure and developer tooling. Requires 8+ years of production software experience, strong backend skills, and expertise with LLM evaluation and agentic systems.