# Software Engineer, ML Infrastructure Platform

**Company:** [Nuro](https://hotfix.jobs/companies/nuro)
**Location:** Mountain View, CA
**Role:** ML Engineering
**Salary:** $160k – $241k/yr
**Experience:** 1+ years
**Skills:** Python, C++, Go, Kubernetes, Distributed Systems, GCP, nccl, gpu training, Data Pipelines, streaming data, ml orchestration, Observability, Reinforcement Learning, batch processing, Incident Response
**Posted:** 2026-08-11

> Build and operate the infrastructure powering large-scale machine-learning training for autonomous-driving systems. The role requires Python proficiency, Kubernetes production experience, distributed-systems expertise, and ownership of reliability, observability, and operational maturity.

## Job Description

## Responsibilities

- Contribute to Nuro’s training infrastructure across multi-generation accelerators and multi-cluster scheduling and orchestration.
- Design and operate large-scale data pipelines, including batch and streaming ingestion, storage layouts, and high-throughput data generation and storage.
- Design and develop agentic-first ML workflows spanning data, training, and evaluation pipelines that are introspectable, reproducible, and easy for autonomy teams to run and extend.
- Own reliability for critical training and release pipelines by instrumenting them, defining meaningful alerting, and building on-call and incident-response practices.

## Requirements

- Bachelor’s, master’s, or doctoral degree in Computer Science, Electrical Engineering, or a closely related field.
- At least 1 year of relevant professional experience.
- Willingness to deep-dive into implementation and raise technical and operational standards.
- Demonstrated ownership mindset, including driving systems toward operational maturity through monitoring, alerting, and runbooks.
- Strong proficiency in Python and comfort with C++, Go, or a similar systems language.
- Hands-on experience running production infrastructure on Kubernetes.
- Solid distributed-systems fundamentals and ability to reason about performance, failure modes, and reliability across complex systems.

## Nice-to-Haves

- Strong working knowledge of Google Cloud.
- Experience building large-scale data-generation pipelines.
- Experience with Kubernetes-native orchestration for ML workloads.
- Knowledge of GPU and distributed-training internals, including NCCL and collective communication.
- Familiarity with GPU and training observability tools and using them to diagnose bottlenecks.
- Track record of reducing infrastructure costs while improving reliability.

## Compensation and Benefits

- Base pay range: **$160,360–$240,540**.
- Eligible for an annual performance bonus, equity, and a competitive benefits package.

## Similar roles

- [Software Engineer, ML Inference Platform](https://hotfix.jobs/jobs/74081e27-c2cd-4485-b56d-22edd2ae9e1d) - Nuro - Mountain View, CA - $160k – $241k/yr
- [Software Engineer, ML Infrastructure, Optimization](https://hotfix.jobs/jobs/a6b5f42b-e6f7-497c-adda-7a35b24c3eb4) - Nuro - Mountain View, CA - $160k – $241k/yr
- [Logistics Research Team](https://hotfix.jobs/jobs/929f6b1a-65f4-42d3-9422-af71df3f30bb) - Sprinter Health - San Francisco, CA - $160k – $200k/yr
- [Software Engineer, Model Performance Tooling](https://hotfix.jobs/jobs/26ede8a1-a8ce-494b-bcb9-87270ba2d333) - Baseten - San Francisco, CA - $160k – $200k/yr
- [Applied Scientist II](https://hotfix.jobs/jobs/c6c7879a-d2d0-4c0c-a46c-43a9930a8457) - Garner Health - New York, NY - $158k – $190k/yr

**Apply:** https://hotfix.jobs/jobs/5ccc0259-2d85-4c10-8ed4-76d8ac33c4ea
**Canonical:** https://hotfix.jobs/jobs/5ccc0259-2d85-4c10-8ed4-76d8ac33c4ea