Software Engineer, ML Performance Optimization
Drive ML performance optimization initiatives to make autonomous driving models faster and more efficient using distributed training, quantization, distillation, and profiling tools.
About the job
Responsibilities
- Design, implement, and operate cutting-edge ML Training OR Inference performance optimization techniques to scale VLM, VLA, and Foundational models and deploy them efficiently in robotaxis.
- Collaborate closely with cross-functional teams, including ML researchers, software engineers, data engineers, and hardware engineers, to define requirements and align on architectural decisions.
Requirements
- 4+ years of total experience, including 2+ years of working on large-scale model training or inference platforms.
- Experience with training frameworks like PyTorch, leveraging GPUs efficiently for distributed model training.
- Experience with GPU-accelerated inference using TensorRT or similar frameworks.
- Experience using profiling tools like NVIDIA's Nsight or PyTorch's Profiler for identifying model training and serving bottlenecks.
- Proficient in Python or C++.
Nice-to-Haves
- Experience with distributed training techniques, quantization, distillation, and pruning.
- Work with SOTA accelerators and inference optimization frameworks.
Skills
PyTorch, TensorRT, Python, C++, Nvidia Nsight, Pytorch Profiler, Distributed Training, Quantization, Distillation, Pruning
Similar jobs
ML Engineering jobsBuild trustworthy infrastructure for production LLM agents, closed-loop evaluation, and autonomous research workflows. The role requires strong Python and distributed-systems experience, hands-on LLM post-training and inference knowledge, and experience operating agent systems at scale.
Build production-grade AI agents, evaluation infrastructure, and developer tooling that make AI-assisted engineering faster, safer, and reusable across teams. The role requires software engineering experience, platform or internal developer-product experience, and hands-on expertise with LLM integration and orchestration.
Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Build and operate production ML infrastructure spanning training, deployment, serving, monitoring, data pipelines, and feedback-driven retraining. The role requires strong MLOps and DevOps experience, Python and SQL proficiency, and ownership of reliable cloud-based systems.
Build and scale post-training, reinforcement-learning, evaluation, and inference systems for long-horizon agents operating over complex enterprise software. The role requires strong Python and PyTorch or JAX skills, distributed GPU experience, empirical rigor, and the ability to take research results into production.