# Senior Software Engineer, ML Infrastructure Platform

**Company:** [Nuro](https://hotfix.jobs/companies/nuro)
**Location:** Mountain View, CA, California
**Role:** ML Engineering
**Salary:** $194k – $291k/yr
**Experience:** 5+ years
**Skills:** Python, C++, Go, Kubernetes, GCP, Distributed Systems, Gpu Training, Nccl, Ml Workflows, Data Pipelines, Streaming Ingestion, Observability, Multi-Cluster Scheduling, Reinforcement Learning
**Posted:** 2026-08-11

> Build and operate large-scale infrastructure for autonomous-driving model training, including distributed GPU systems, data pipelines, ML workflows, and reliability tooling. The role requires 3+ years of experience, strong Python and systems-language skills, Kubernetes expertise, and distributed-systems fundamentals.

## Job Description

## Responsibilities
- Contribute to training infrastructure spanning multi-generation accelerators, multi-cluster scheduling, and orchestration.
- Design and operate large-scale data pipelines, including batch and streaming ingestion, storage layout, and high-throughput data generation and storage.
- Design and develop agentic-first ML workflows covering data-to-training-to-evaluation pipelines that are introspectable, reproducible, and easy for autonomy teams to run and extend.
- Own reliability for critical training and release pipelines by instrumenting them, defining meaningful alerting, and building on-call and incident-response practices.

## Requirements
- Bachelor's, master's, or doctoral degree in Computer Science, Electrical Engineering, or a closely related field.
- 3+ years of relevant professional experience.
- Willingness to deep-dive into implementation and raise technical and operational standards.
- Demonstrated ownership mindset, including driving systems to operational maturity through monitoring, alerting, and runbooks.
- Strong proficiency in Python and comfort with C++, Go, or a similar systems language.
- Hands-on experience running production infrastructure on Kubernetes.
- Solid distributed-systems fundamentals and ability to reason about performance, failure modes, and reliability across complex systems.

## Nice-to-Haves
- Strong working knowledge of Google Cloud.
- Experience building large-scale data-generation pipelines.
- Experience with Kubernetes-native orchestration for ML workloads.
- Knowledge of GPU and distributed-training internals, including NCCL and collective communication.
- Familiarity with GPU and training observability tooling and using it to diagnose bottlenecks.
- Track record of reducing infrastructure costs while improving reliability.

## Compensation and Benefits
- Base pay range: **$193,930–$291,150**.
- Eligible for an annual performance bonus, equity, and competitive benefits.

## Similar jobs

- [Senior ML/AI Modeler, Risk Automation Machine Learning](https://hotfix.jobs/jobs/303b4f2e-9bc8-4034-8fa1-d17c71bf6f1e) - Square - Remote - $195k – $343k/yr
- [Senior Software Engineer, Autonomy Capabilities](https://hotfix.jobs/jobs/ae9ec2ef-285e-42da-9740-0b3ce5faf8cf) - Shield AI - San Mateo, CA - $196k – $294k/yr
- [Senior Machine Learning Engineer, Payments](https://hotfix.jobs/jobs/4dcd0258-2a7a-4e22-a27e-f2e1426de86e) - Airbnb - Remote - $191k – $223k/yr
- [Senior Machine Learning Infrastructure Engineer, Embedding Platform](https://hotfix.jobs/jobs/f7e37122-fd13-4390-b7b8-27c86092f1ae) - Reddit - Remote - $191k – $267k/yr
- [Senior Applied AI Engineer](https://hotfix.jobs/jobs/03330d35-9411-4346-b3db-57b680d941c7) - Mintlify - San Francisco, CA - $190k – $265k/yr

**Apply:** https://hotfix.jobs/jobs/42001ce3-19de-49b5-8883-d001ca38069f
**Canonical:** https://hotfix.jobs/jobs/42001ce3-19de-49b5-8883-d001ca38069f