Skip to content
NuroNuro

Software Engineer, ML Infrastructure

Build and scale ML infrastructure platform for autonomous vehicle development, focusing on automated resource provisioning, high-performance workload scheduling, and petabyte-scale data processing pipelines.

About the job

Responsibilities

  • Build and evolve the core ML infrastructure platform providing researchers and engineers seamless access to compute and data resources
  • Scale automated Infrastructure-as-Code (IaC) pipelines to manage thousands of GPU/CPU nodes across diverse environments
  • Design and optimize workload orchestration to maximize hardware utilization, minimize job wait times, and handle massive-scale distributed training
  • Design robust pipelines for extraction and transformation of petabyte-scale sensor and telemetry data into ML-ready formats
  • Implement robust feature caching and storage solutions to reduce redundant computations and ensure low-latency access to pre-computed features
  • Contribute to a unified ML platform that abstracts complex cloud infrastructure for end-users

Requirements

  • 3+ years of professional experience in ML Infrastructure, Backend Platform Engineering, or Distributed Systems
  • Deep familiarity with modern Infrastructure-as-Code and provisioning tools such as Terraform, Pulumi, or Crossplane
  • Hands-on experience building or managing large-scale orchestrators for compute-heavy workloads (e.g., Kubernetes, KubeRay, Ray, Slurm, or Volcano)
  • Proficiency in at least one distributed processing framework, such as Apache Spark or Apache Beam, for large-scale data extraction and transformation
  • Experience implementing or maintaining feature stores and caching layers (e.g., Feast, Hopsworks, or Redis-based custom caching)
  • Strong understanding of distributed systems, networking, and storage bottlenecks in the context of high-performance computing

Nice-to-Haves

  • Active contributor to open-source projects in the MLOps or Cloud-Native ecosystem (e.g., CNCF, Ray, or Kubeflow communities)
  • Experience with high-performance storage systems (e.g., Lustre, Ceph, or specialized NVMe caching) for ML data loading
  • Knowledge of cost-optimization strategies for large-scale GPU clusters in public clouds (AWS, GCP, or Azure)

Skills

Terraform, Pulumi, Crossplane, Kubernetes, Kuberay, Ray, Slurm, Volcano, Spark, Apache Beam, Feast, Hopsworks, Redis, Infrastructure As Code, Distributed Systems

Pindrop

Pindrop

United States

Research Scientist II
$160k+/yrRemote3+ YOEML Engineering

Research Scientist II building and improving fraud risk models and scam detection systems using audio, behavioral, and metadata signals. Requires an advanced degree and 3+ years of applied ML experience with Python and modern ML frameworks.

Office Hours

Office Hours

San Francisco, CA
Software Engineer - Benchmarking
$160k+/yrRemote4+ YOEML Engineering

Build reproducible systems for AI model benchmarking, including datasets, evaluation pipelines, containerized environments, scoreboards, and analysis tools. The role requires 4+ years of professional engineering experience, strong Python, dataset rigor, and Docker expertise.

Applied Intuition

Applied Intuition

Sunnyvale, CA

Software Engineer - Prediction and Planning ML
$151k+/yrOn-site3+ YOEML Engineering

Develop and deploy ML-first behavior prediction and planning systems for autonomous vehicles, forecasting the motion and interactions of road users. Requires a bachelor's degree, deep learning lifecycle expertise, and at least three years of production software experience with C++ or Python.

Clay

Clay

New York, NY

Software Engineer, Applied AI
$170k+/yrHybridML Engineering

Build and ship production AI agents and the platform infrastructure that makes them reliable, steerable, and measurable. The role requires strong backend fundamentals, production LLM or agent experience, and expertise in evaluations, retrieval, orchestration, or tool-use design.

Clay

Clay

San Francisco, CA

Machine Learning Engineer
$170k+/yrHybrid5+ YOEML Engineering

Build and ship production machine-learning systems that learn from customer data and behavior, including recommendations, LLM-powered features, evaluation systems, and ML infrastructure. The role requires 5+ years of ML engineering or ML-heavy software engineering experience and strong production systems expertise.