Skip to content
PerplexityPerplexitySan Francisco, CA

Member of Technical Staff

Build and operate a self-serve GPU platform for training and inference workloads across multiple clouds and providers. Requires deep Kubernetes and distributed systems expertise plus experience managing large-scale GPU fleets.

250k – 485k/yr
Remote7+ YOEDevOps / SRE

About the role

Responsibilities

  • Build a self-serve compute platform that lets inference engineers and researchers launch training jobs and operate inference services without managing GPU provisioning, cluster configuration, or provider-specific infrastructure.
  • Operate the GPU fleet, including provisioning, lifecycle management, reliability, and capacity integration across providers.
  • Solve for GPU scarcity by building scheduling and placement logic that finds available capacity across providers, packs it efficiently, and gets the right workload onto the right hardware.
  • Support long-running distributed training jobs while guaranteeing the availability and latency of production inference services on the same fleet.
  • Own Kubernetes for GPU orchestration, including writing operators and CRDs, and managing many clusters across providers.
  • Build fault tolerance, autoscaling, and observability to keep the fleet utilized and let workloads survive node loss, provider hiccups, and capacity shifts without human intervention.
  • Set technical direction across teams by partnering with inference and cloud infrastructure engineers to turn operational constraints into a coherent platform architecture and roadmap.

Qualifications

  • Deep Kubernetes experience with custom operators, CRDs, and multi-cluster federation.
  • Experience managing GPU clusters at scale, including NVIDIA hardware, CUDA, and high-performance networking (InfiniBand or RoCE).
  • Experience orchestrating compute across multiple clouds (CoreWeave, AWS, GCP, or similar).
  • Strong distributed systems fundamentals in scheduling, resource allocation, and fault tolerance under load.
  • Proficiency writing infrastructure and systems-level code in Go, Rust, or C++.
  • Experience supporting both long-running training jobs and high-availability inference services.

Nice-to-Haves

  • Experience with inference serving stacks: vLLM, SGLang, or TensorRT-LLM.
  • Experience with Slurm or other HPC schedulers.
  • GPU kernel work in CUDA or Triton.
  • Experience with high-speed interconnects: InfiniBand, RoCE, or RDMA in production.
  • Observability for ML workloads: Prometheus, Grafana, or Weights & Biases.

Skills

Kubernetesgpu clustersCUDAInfiniBandroceGoRustC++Distributed Systemsmulti-cloudvLLMtensorrt-llmPrometheusGrafana

Similar roles

DevOps / SRE jobs
Perplexity

Member of Technical Staff

PerplexitySan Francisco, CA +1

Hands-on technical role building AI-powered tools, infrastructure, and processes to accelerate engineering velocity and product delivery at an AI search company.

250k – 405k/yr
Hybrid5+ YOEDevOps / SRE
Together AI

Staff Engineer, Distributed Storage and HPC & AI Infrastructure

Together AISan Francisco, CA

Design and operate multi-petabyte distributed storage systems for large-scale AI training and inference, integrating parallel filesystems and building Kubernetes-native storage platforms.

250k – 300k/yr
On-site8+ YOEDevOps / SRE
Zoox

Staff Site Reliability Engineer

ZooxFoster City, CA

Zoox is seeking a Staff Site Reliability Engineer to lead source control, owning the technical strategy and roadmap for their Git-based monorepo. This role involves migrating from GitHub Enterprise to GitHub Cloud, building developer tooling, and partnering with various teams to enhance source control as a strategic asset.

250k – 300k/yr
Hybrid7+ YOEDevOps / SRE
Ambience Healthcare

Staff Engineer, Agent Productivity

Ambience HealthcareSan Francisco, CA

Owns developer productivity infrastructure including dev environments, CI workflows, verification systems, internal tooling, and secure agent tool access to enable efficient engineering and AI agent development. Requires 7+ years experience with strong backend skills in TypeScript, Rust, Go, or Python.

250k – 300k/yr
Hybrid7+ YOEDevOps / SRE
Envoy

Member of Technical Staff, Infrastructure

EnvoySan Francisco, CA

Senior infrastructure engineer owning core platform reliability, evolving Kubernetes and AWS systems with Terraform, defining SLOs, and reducing operational toil through automation and paved paths. Requires 10+ years experience and strong software engineering skills.

250k – 280k/yr
On-site10+ YOEDevOps / SRE