Staff Software Engineer, ML Performance Optimization
Leads ML performance optimization for training and inference platforms in autonomous driving, collaborating across teams to enhance efficiency using PyTorch, TensorRT, and profiling tools. Requires strong Python/C++ skills, GPU expertise, and 4+ years in large-scale ML platforms; leads engineering team.
About the job
Responsibilities
- Develop and execute a strategic vision for the ML Performance Optimization team to unlock ML innovation in autonomous driving and rider experience.
- Lead the design, implementation, and operation of cutting-edge ML Training and inference performance optimization techniques.
- Collaborate closely with cross-functional teams, including ML researchers, software engineers, data engineers, and hardware engineers, to define requirements and align on architectural decisions.
- Enable the engineers in the team to grow their careers by providing technical guidance and mentorship.
Qualifications
- Strong experience with training frameworks like PyTorch, leveraging GPUs efficiently for distributed model training.
- Experience with GPU-accelerated inference using TensorRT, Ray Serve, or similar frameworks.
- Experience using profiling tools like NVIDIA Nsight or PyTorch Profiler for identifying model training and serving bottlenecks.
- Proficient in Python and C++.
- Experience with model compression techniques to reduce model size and improve performance.
Bonus Qualifications
- 10+ years of total experience, including 4+ years of working on large-scale model training or inference platforms.
- Excellent leadership skills with a demonstrated ability to lead high-performing engineering teams.
Skills
PyTorch, TensorRT, Ray Serve, Nvidia Nsight, Pytorch Profiler, Python, C++, GPU, Distributed Training, Model Compression
Similar jobs
ML Engineering jobsLeads the roadmap and technical vision for Snowflake Feature Store, building reliable, high-performance machine learning platform capabilities and supporting technical execution across partner teams. Requires 10+ years of experience with data-serving infrastructure or ML platforms, plus Java and Python expertise.
Senior Staff ML Engineer fine-tunes and optimizes state-of-the-art LLMs for Airbnb's customer support AI products, including AI assistants and autonomous agents. Partners cross-functionally to productionize models at scale. Requires PhD and 10+ years experience with PyTorch.
Leads the design, implementation, integration, and field validation of tactical autonomy and multi-agent coordination capabilities for unmanned platforms. Requires 7+ years of relevant experience, production C++, technical leadership, and eligibility for a U.S. Secret clearance.
Staff Machine Learning Engineer building and operating production ML systems for causal marketing measurement, optimization, and planning. The role requires deep statistical and machine learning expertise, production programming experience, cross-functional collaboration, and technical mentorship.
Staff engineer responsible for designing and scaling the infrastructure, execution environments, verifiers, and tooling used to train and evaluate AI agents. Requires 8+ years of software engineering experience, strong Python and distributed-systems expertise, and familiarity with sandboxing, high-throughput systems, and LLM workflows.