Build and operate the infrastructure powering large-scale machine-learning training for autonomous-driving systems. The role requires Python proficiency, Kubernetes production experience, distributed-systems expertise, and ownership of reliability, observability, and operational maturity.
160k – 241k/yr
On-site1+ YOEML Engineering
About the role
Responsibilities
Contribute to Nuro’s training infrastructure across multi-generation accelerators and multi-cluster scheduling and orchestration.
Design and operate large-scale data pipelines, including batch and streaming ingestion, storage layouts, and high-throughput data generation and storage.
Design and develop agentic-first ML workflows spanning data, training, and evaluation pipelines that are introspectable, reproducible, and easy for autonomy teams to run and extend.
Own reliability for critical training and release pipelines by instrumenting them, defining meaningful alerting, and building on-call and incident-response practices.
Requirements
Bachelor’s, master’s, or doctoral degree in Computer Science, Electrical Engineering, or a closely related field.
At least 1 year of relevant professional experience.
Willingness to deep-dive into implementation and raise technical and operational standards.
Demonstrated ownership mindset, including driving systems toward operational maturity through monitoring, alerting, and runbooks.
Strong proficiency in Python and comfort with C++, Go, or a similar systems language.
Hands-on experience running production infrastructure on Kubernetes.
Solid distributed-systems fundamentals and ability to reason about performance, failure modes, and reliability across complex systems.
Nice-to-Haves
Strong working knowledge of Google Cloud.
Experience building large-scale data-generation pipelines.
Experience with Kubernetes-native orchestration for ML workloads.
Knowledge of GPU and distributed-training internals, including NCCL and collective communication.
Familiarity with GPU and training observability tools and using them to diagnose bottlenecks.
Track record of reducing infrastructure costs while improving reliability.
Compensation and Benefits
Base pay range: $160,360–$240,540.
Eligible for an annual performance bonus, equity, and a competitive benefits package.
Build and maintain machine learning infrastructure for autonomy teams, including model pipelines, observability, inference serving, and compiler platforms. The role requires a relevant degree, at least one year of experience, strong Python skills, and familiarity with C++.
160k – 241k/yrOn-site1+ YOEML Engineering
Software Engineer, ML Infrastructure, Optimization
NuroMountain View, CA
Build and optimize ML infrastructure for autonomous vehicles, focusing on model optimization, compilers, and deployment across the autonomy stack. Requires 2+ years in ML optimization and strong Python/C++/CUDA skills.
160k – 241k/yrOn-site2+ YOEML Engineering
Logistics Research Team
Sprinter HealthSan Francisco, CA
Software Engineer on the Logistics Optimization team designing and implementing algorithms for clinician routing, scheduling, dispatch, simulations, and predictive models to optimize in-home healthcare delivery at national scale. Requires 2-3 years software engineering experience with optimization, forecasting or simulation systems, preferably in TypeScript/Python.
160k – 200k/yrHybrid2+ YOEML Engineering
Software Engineer, Model Performance Tooling
BasetenSan Francisco, CA
Builds performance benchmarking, diagnostic, and optimization tools for LLM inference on GPU clusters. Early-career role requiring Python proficiency, systems curiosity, and interest in AI hardware—no prior experience needed.
160k – 200k/yrOn-siteEntry levelML Engineering
Applied Scientist II
Garner HealthNew York, NY
Build and ship production algorithmic systems that improve healthcare quality, access, and cost outcomes. The role combines machine learning, optimization, experimentation, and LLM productionization, requiring at least two years of relevant industry or advanced-degree experience.