# Member of Technical Staff

**Company:** [Perplexity](https://hotfix.jobs/companies/perplexity)
**Location:** Remote
**Role:** DevOps / SRE
**Salary:** $250k – $485k/yr
**Experience:** 7+ years
**Skills:** Kubernetes, gpu clusters, CUDA, InfiniBand, roce, Go, Rust, C++, Distributed Systems, multi-cloud, vLLM, tensorrt-llm, Prometheus, Grafana
**Posted:** 2026-07-16

> Build and operate a self-serve GPU platform for training and inference workloads across multiple clouds and providers. Requires deep Kubernetes and distributed systems expertise plus experience managing large-scale GPU fleets.

## Job Description

## Responsibilities
- Build a self-serve compute platform that lets inference engineers and researchers launch training jobs and operate inference services without managing GPU provisioning, cluster configuration, or provider-specific infrastructure.
- Operate the GPU fleet, including provisioning, lifecycle management, reliability, and capacity integration across providers.
- Solve for GPU scarcity by building scheduling and placement logic that finds available capacity across providers, packs it efficiently, and gets the right workload onto the right hardware.
- Support long-running distributed training jobs while guaranteeing the availability and latency of production inference services on the same fleet.
- Own Kubernetes for GPU orchestration, including writing operators and CRDs, and managing many clusters across providers.
- Build fault tolerance, autoscaling, and observability to keep the fleet utilized and let workloads survive node loss, provider hiccups, and capacity shifts without human intervention.
- Set technical direction across teams by partnering with inference and cloud infrastructure engineers to turn operational constraints into a coherent platform architecture and roadmap.

## Qualifications
- Deep Kubernetes experience with custom operators, CRDs, and multi-cluster federation.
- Experience managing GPU clusters at scale, including NVIDIA hardware, CUDA, and high-performance networking (InfiniBand or RoCE).
- Experience orchestrating compute across multiple clouds (CoreWeave, AWS, GCP, or similar).
- Strong distributed systems fundamentals in scheduling, resource allocation, and fault tolerance under load.
- Proficiency writing infrastructure and systems-level code in Go, Rust, or C++.
- Experience supporting both long-running training jobs and high-availability inference services.

## Nice-to-Haves
- Experience with inference serving stacks: vLLM, SGLang, or TensorRT-LLM.
- Experience with Slurm or other HPC schedulers.
- GPU kernel work in CUDA or Triton.
- Experience with high-speed interconnects: InfiniBand, RoCE, or RDMA in production.
- Observability for ML workloads: Prometheus, Grafana, or Weights & Biases.

## Similar roles

- [Member of Technical Staff](https://hotfix.jobs/jobs/ca5edb5f-58ba-4e0c-a075-25c3fcc61007) - Perplexity - San Francisco, CA - $250k – $405k/yr
- [Staff Engineer, Distributed Storage and HPC & AI Infrastructure](https://hotfix.jobs/jobs/d0fb38e7-169a-4b2d-abb6-6f845d3f381f) - Together AI - San Francisco, CA - $250k – $300k/yr
- [Staff Site Reliability Engineer](https://hotfix.jobs/jobs/e2420bd8-af00-4aa8-9bcb-34da7d97403d) - Zoox - Foster City, CA - $250k – $300k/yr
- [Staff Engineer, Agent Productivity](https://hotfix.jobs/jobs/b860cf40-a9f4-43ad-961a-fc3aaf976c3d) - Ambience Healthcare - San Francisco, CA - $250k – $300k/yr
- [Member of Technical Staff, Infrastructure](https://hotfix.jobs/jobs/0b3aa2e6-0511-4a76-b1d9-f69a3738db4d) - Envoy - San Francisco, CA - $250k – $280k/yr

**Apply:** https://hotfix.jobs/jobs/996eec50-5e6f-42cb-97a1-0219ffb52860
**Canonical:** https://hotfix.jobs/jobs/996eec50-5e6f-42cb-97a1-0219ffb52860