Skip to content
PerplexityPerplexitySan Francisco, CA

Member of Technical Staff

Owns a multi-cloud GPU infrastructure platform that enables training and inference workloads through self-service orchestration. The role requires deep Kubernetes, distributed systems, GPU networking, and systems programming experience.

250k – 485k/yr
RemoteDevOps / SRE

About the role

Responsibilities

  • Build a self-serve compute platform for launching training jobs and operating inference services without managing GPU provisioning, cluster configuration, or provider-specific infrastructure.
  • Own GPU fleet provisioning, lifecycle management, reliability, and capacity integration across cloud providers.
  • Build scheduling and placement logic to find available capacity, efficiently pack workloads, and match workloads to appropriate hardware.
  • Support long-running distributed training jobs and highly available, low-latency production inference services on the same fleet.
  • Write Kubernetes operators and CRDs and manage multi-provider clusters.
  • Build fault tolerance, autoscaling, and observability for resilience against node loss, provider failures, and capacity shifts.
  • Set technical direction and partner with inference and cloud infrastructure engineers on platform architecture and roadmap.

Requirements

  • Deep Kubernetes experience, including custom operators, CRDs, and multi-cluster federation.
  • Experience managing GPU clusters at scale with NVIDIA hardware, CUDA, and high-performance networking such as InfiniBand or RoCE.
  • Experience orchestrating compute across multiple clouds, such as CoreWeave, AWS, or GCP.
  • Strong distributed-systems fundamentals, including scheduling, resource allocation, and fault tolerance under load.
  • Infrastructure and systems-level programming experience in Go, Rust, or C++.
  • Experience supporting long-running training jobs and high-availability inference services.
  • End-to-end ownership and comfort working independently in ambiguous environments.

Nice to Have

  • Experience with inference-serving stacks such as vLLM, SGLang, or TensorRT-LLM.
  • Experience with Slurm or other HPC schedulers.
  • GPU kernel experience with CUDA or Triton.
  • Production experience with InfiniBand, RoCE, or RDMA.
  • Observability experience for ML workloads with Prometheus, Grafana, or Weights & Biases.

Skills

KubernetesGoRustC++CUDAnvidia gpusInfiniBandrocerdmaAWSGCPcoreweavevLLMsglangtensorrt-llm

Similar roles

DevOps / SRE jobs
Crusoe

Senior Staff Deployment Automation Engineer

CrusoeSan Francisco, CA +2

Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.

250k – 300k/yrOn-site12+ YOEDevOps / SRE
Crusoe

Senior Staff Software Engineer, DC Infrastructure

CrusoeSan Francisco, CA +1

Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.

250k – 300k/yrOn-site7+ YOEDevOps / SRE
Perplexity

Member of Technical Staff

PerplexitySan Francisco, CA +1

Hands-on technical role building AI-powered tools, infrastructure, and processes to accelerate engineering velocity and product delivery at an AI search company.

250k – 405k/yrHybrid5+ YOEDevOps / SRE
Together AI

Staff Engineer, Distributed Storage and HPC & AI Infrastructure

Together AISan Francisco, CA

Design and operate multi-petabyte distributed storage systems for large-scale AI training and inference, integrating parallel filesystems and building Kubernetes-native storage platforms.

250k – 300k/yrOn-site8+ YOEDevOps / SRE
Zoox

Staff Site Reliability Engineer

ZooxFoster City, CA

Zoox is seeking a Staff Site Reliability Engineer to lead source control, owning the technical strategy and roadmap for their Git-based monorepo. This role involves migrating from GitHub Enterprise to GitHub Cloud, building developer tooling, and partnering with various teams to enhance source control as a strategic asset.

250k – 300k/yrHybrid7+ YOEDevOps / SRE