Owns a multi-cloud GPU infrastructure platform that enables training and inference workloads through self-service orchestration. The role requires deep Kubernetes, distributed systems, GPU networking, and systems programming experience.
250k – 485k/yr
RemoteDevOps / SRE
About the role
Responsibilities
Build a self-serve compute platform for launching training jobs and operating inference services without managing GPU provisioning, cluster configuration, or provider-specific infrastructure.
Own GPU fleet provisioning, lifecycle management, reliability, and capacity integration across cloud providers.
Build scheduling and placement logic to find available capacity, efficiently pack workloads, and match workloads to appropriate hardware.
Support long-running distributed training jobs and highly available, low-latency production inference services on the same fleet.
Write Kubernetes operators and CRDs and manage multi-provider clusters.
Build fault tolerance, autoscaling, and observability for resilience against node loss, provider failures, and capacity shifts.
Set technical direction and partner with inference and cloud infrastructure engineers on platform architecture and roadmap.
Requirements
Deep Kubernetes experience, including custom operators, CRDs, and multi-cluster federation.
Experience managing GPU clusters at scale with NVIDIA hardware, CUDA, and high-performance networking such as InfiniBand or RoCE.
Experience orchestrating compute across multiple clouds, such as CoreWeave, AWS, or GCP.
Strong distributed-systems fundamentals, including scheduling, resource allocation, and fault tolerance under load.
Infrastructure and systems-level programming experience in Go, Rust, or C++.
Experience supporting long-running training jobs and high-availability inference services.
End-to-end ownership and comfort working independently in ambiguous environments.
Nice to Have
Experience with inference-serving stacks such as vLLM, SGLang, or TensorRT-LLM.
Experience with Slurm or other HPC schedulers.
GPU kernel experience with CUDA or Triton.
Production experience with InfiniBand, RoCE, or RDMA.
Observability experience for ML workloads with Prometheus, Grafana, or Weights & Biases.
Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.
250k – 300k/yrOn-site12+ YOEDevOps / SRE
Senior Staff Software Engineer, DC Infrastructure
CrusoeSan Francisco, CA +1
Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.
250k – 300k/yrOn-site7+ YOEDevOps / SRE
Member of Technical Staff
PerplexitySan Francisco, CA +1
Hands-on technical role building AI-powered tools, infrastructure, and processes to accelerate engineering velocity and product delivery at an AI search company.
250k – 405k/yrHybrid5+ YOEDevOps / SRE
Staff Engineer, Distributed Storage and HPC & AI Infrastructure
Together AISan Francisco, CA
Design and operate multi-petabyte distributed storage systems for large-scale AI training and inference, integrating parallel filesystems and building Kubernetes-native storage platforms.
250k – 300k/yrOn-site8+ YOEDevOps / SRE
Staff Site Reliability Engineer
ZooxFoster City, CA
Zoox is seeking a Staff Site Reliability Engineer to lead source control, owning the technical strategy and roadmap for their Git-based monorepo. This role involves migrating from GitHub Enterprise to GitHub Cloud, building developer tooling, and partnering with various teams to enhance source control as a strategic asset.