Senior Software Engineer, ML Infrastructure Platform
Build and operate large-scale infrastructure for autonomous-driving model training, including distributed GPU systems, data pipelines, ML workflows, and reliability tooling. The role requires 3+ years of experience, strong Python and systems-language skills, Kubernetes expertise, and distributed-systems fundamentals.
About the job
Responsibilities
- Contribute to training infrastructure spanning multi-generation accelerators, multi-cluster scheduling, and orchestration.
- Design and operate large-scale data pipelines, including batch and streaming ingestion, storage layout, and high-throughput data generation and storage.
- Design and develop agentic-first ML workflows covering data-to-training-to-evaluation pipelines that are introspectable, reproducible, and easy for autonomy teams to run and extend.
- Own reliability for critical training and release pipelines by instrumenting them, defining meaningful alerting, and building on-call and incident-response practices.
Requirements
- Bachelor's, master's, or doctoral degree in Computer Science, Electrical Engineering, or a closely related field.
- 3+ years of relevant professional experience.
- Willingness to deep-dive into implementation and raise technical and operational standards.
- Demonstrated ownership mindset, including driving systems to operational maturity through monitoring, alerting, and runbooks.
- Strong proficiency in Python and comfort with C++, Go, or a similar systems language.
- Hands-on experience running production infrastructure on Kubernetes.
- Solid distributed-systems fundamentals and ability to reason about performance, failure modes, and reliability across complex systems.
Nice-to-Haves
- Strong working knowledge of Google Cloud.
- Experience building large-scale data-generation pipelines.
- Experience with Kubernetes-native orchestration for ML workloads.
- Knowledge of GPU and distributed-training internals, including NCCL and collective communication.
- Familiarity with GPU and training observability tooling and using it to diagnose bottlenecks.
- Track record of reducing infrastructure costs while improving reliability.
Compensation and Benefits
- Base pay range: $193,930–$291,150.
- Eligible for an annual performance bonus, equity, and competitive benefits.
Skills
Python, C++, Go, Kubernetes, GCP, Distributed Systems, Gpu Training, Nccl, Ml Workflows, Data Pipelines, Streaming Ingestion, Observability, Multi-Cluster Scheduling, Reinforcement Learning
Similar jobs
ML Engineering jobsLeads the design, deployment, and optimization of agentic and generative AI systems that automate risk and compliance investigations at scale. Requires 8+ years of machine learning modeling experience, production ML expertise, and advanced technical education.
Design and deploy tactical autonomy algorithms and high-performance software for unmanned systems operating in complex, contested environments. The role requires 5+ years of related experience, strong C++ and Python skills, robotics expertise, and the ability to obtain a SECRET clearance.
Senior machine learning engineer who will build and operate large-scale AI systems for Airbnb’s payments ecosystem, including LLM agents, fraud defenses, and personalization. The role requires 5+ years of applied AI/ML experience, strong Python or Java skills, and production MLOps expertise.
Build and operate large-scale machine learning infrastructure and models for Reddit’s recommendation and personalization systems. The role requires 5+ years of ML engineering experience, expertise in deep learning and distributed systems, and proficiency with Python and modern ML frameworks.
Build and lead production AI products, including agent systems that use tools, retrieve context, and complete complex tasks reliably. The role requires strong full-stack engineering, deep language-model experience, product judgment, and ownership from experimentation through production.