Skip to content
NuroNuroMountain View, CA

Software Engineer, ML Infrastructure Platform

Build and operate the infrastructure powering large-scale machine-learning training for autonomous-driving systems. The role requires Python proficiency, Kubernetes production experience, distributed-systems expertise, and ownership of reliability, observability, and operational maturity.

160k – 241k/yr
On-site1+ YOEML Engineering

About the role

Responsibilities

  • Contribute to Nuro’s training infrastructure across multi-generation accelerators and multi-cluster scheduling and orchestration.
  • Design and operate large-scale data pipelines, including batch and streaming ingestion, storage layouts, and high-throughput data generation and storage.
  • Design and develop agentic-first ML workflows spanning data, training, and evaluation pipelines that are introspectable, reproducible, and easy for autonomy teams to run and extend.
  • Own reliability for critical training and release pipelines by instrumenting them, defining meaningful alerting, and building on-call and incident-response practices.

Requirements

  • Bachelor’s, master’s, or doctoral degree in Computer Science, Electrical Engineering, or a closely related field.
  • At least 1 year of relevant professional experience.
  • Willingness to deep-dive into implementation and raise technical and operational standards.
  • Demonstrated ownership mindset, including driving systems toward operational maturity through monitoring, alerting, and runbooks.
  • Strong proficiency in Python and comfort with C++, Go, or a similar systems language.
  • Hands-on experience running production infrastructure on Kubernetes.
  • Solid distributed-systems fundamentals and ability to reason about performance, failure modes, and reliability across complex systems.

Nice-to-Haves

  • Strong working knowledge of Google Cloud.
  • Experience building large-scale data-generation pipelines.
  • Experience with Kubernetes-native orchestration for ML workloads.
  • Knowledge of GPU and distributed-training internals, including NCCL and collective communication.
  • Familiarity with GPU and training observability tools and using them to diagnose bottlenecks.
  • Track record of reducing infrastructure costs while improving reliability.

Compensation and Benefits

  • Base pay range: $160,360–$240,540.
  • Eligible for an annual performance bonus, equity, and a competitive benefits package.

Skills

PythonC++GoKubernetesDistributed SystemsGCPncclgpu trainingData Pipelinesstreaming dataml orchestrationObservabilityReinforcement Learningbatch processingIncident Response

Similar roles

ML Engineering jobs
Nuro

Software Engineer, ML Inference Platform

NuroMountain View, CA

Build and maintain machine learning infrastructure for autonomy teams, including model pipelines, observability, inference serving, and compiler platforms. The role requires a relevant degree, at least one year of experience, strong Python skills, and familiarity with C++.

160k – 241k/yrOn-site1+ YOEML Engineering
Nuro

Software Engineer, ML Infrastructure, Optimization

NuroMountain View, CA

Build and optimize ML infrastructure for autonomous vehicles, focusing on model optimization, compilers, and deployment across the autonomy stack. Requires 2+ years in ML optimization and strong Python/C++/CUDA skills.

160k – 241k/yrOn-site2+ YOEML Engineering
Sprinter Health

Logistics Research Team

Sprinter HealthSan Francisco, CA

Software Engineer on the Logistics Optimization team designing and implementing algorithms for clinician routing, scheduling, dispatch, simulations, and predictive models to optimize in-home healthcare delivery at national scale. Requires 2-3 years software engineering experience with optimization, forecasting or simulation systems, preferably in TypeScript/Python.

160k – 200k/yrHybrid2+ YOEML Engineering
Baseten

Software Engineer, Model Performance Tooling

BasetenSan Francisco, CA

Builds performance benchmarking, diagnostic, and optimization tools for LLM inference on GPU clusters. Early-career role requiring Python proficiency, systems curiosity, and interest in AI hardware—no prior experience needed.

160k – 200k/yrOn-siteEntry levelML Engineering
Garner Health

Applied Scientist II

Garner HealthNew York, NY

Build and ship production algorithmic systems that improve healthcare quality, access, and cost outcomes. The role combines machine learning, optimization, experimentation, and LLM productionization, requiring at least two years of relevant industry or advanced-degree experience.

158k – 190k/yrHybrid2+ YOEML Engineering