Senior Software Engineer, ML Infrastructure
Build and scale ML infrastructure platform for autonomous vehicle model development, focusing on automated resource provisioning, high-performance workload scheduling, and petabyte-scale data processing pipelines.
About the job
Responsibilities
- Build and evolve the core ML infrastructure platform providing researchers and engineers seamless access to compute and data resources
- Scale automated Infrastructure-as-Code (IaC) pipelines to manage thousands of GPU/CPU nodes across diverse environments
- Design and optimize workload orchestration to maximize hardware utilization, minimize job wait times, and handle massive-scale distributed training
- Design robust pipelines for extraction and transformation of petabyte-scale sensor and telemetry data into ML-ready formats
- Implement robust feature caching and storage solutions to reduce redundant computations and ensure low-latency access to pre-computed features
- Contribute to a unified ML platform that abstracts complex cloud infrastructure for end-users
Requirements
- 4+ years of professional experience in ML Infrastructure, Backend Platform Engineering, or Distributed Systems
- Deep familiarity with modern Infrastructure-as-Code and provisioning tools such as Terraform, Pulumi, or Crossplane
- Hands-on experience building or managing large-scale orchestrators for compute-heavy workloads (e.g., Kubernetes, KubeRay, Ray, Slurm, or Volcano)
- Proficiency in at least one distributed processing framework, such as Apache Spark or Apache Beam, for large-scale data extraction and transformation
- Experience implementing or maintaining feature stores and caching layers (e.g., Feast, Hopsworks, or Redis-based custom caching)
- Strong understanding of distributed systems, networking, and storage bottlenecks in the context of high-performance computing
Nice-to-Haves
- Active contributor to open-source projects in the MLOps or Cloud-Native ecosystem (e.g., CNCF, Ray, or Kubeflow communities)
- Experience with high-performance storage systems (e.g., Lustre, Ceph, or specialized NVMe caching) for ML data loading
- Knowledge of cost-optimization strategies for large-scale GPU clusters in public clouds (AWS, GCP, or Azure)
Skills
Terraform, Pulumi, Crossplane, Kubernetes, Ray, Slurm, Spark, Apache Beam, Feast, Hopsworks, Redis, AWS, GCP, Azure
Similar jobs
ML Engineering jobsBuild and operate large-scale infrastructure for autonomous-driving model training, including distributed GPU systems, data pipelines, ML workflows, and reliability tooling. The role requires 3+ years of experience, strong Python and systems-language skills, Kubernetes expertise, and distributed-systems fundamentals.
Leads the design, deployment, and optimization of agentic and generative AI systems that automate risk and compliance investigations at scale. Requires 8+ years of machine learning modeling experience, production ML expertise, and advanced technical education.
Design and deploy tactical autonomy algorithms and high-performance software for unmanned systems operating in complex, contested environments. The role requires 5+ years of related experience, strong C++ and Python skills, robotics expertise, and the ability to obtain a SECRET clearance.
Senior machine learning engineer who will build and operate large-scale AI systems for Airbnb’s payments ecosystem, including LLM agents, fraud defenses, and personalization. The role requires 5+ years of applied AI/ML experience, strong Python or Java skills, and production MLOps expertise.
Build and operate large-scale machine learning infrastructure and models for Reddit’s recommendation and personalization systems. The role requires 5+ years of ML engineering experience, expertise in deep learning and distributed systems, and proficiency with Python and modern ML frameworks.