# Member of Technical Staff

**Company:** [Perplexity](https://hotfix.jobs/companies/perplexity)
**Location:** Remote
**Role:** DevOps / SRE
**Salary:** $250k – $485k/yr
**Skills:** Kubernetes, Go, Rust, C++, CUDA, nvidia gpus, InfiniBand, roce, rdma, AWS, GCP, coreweave, vLLM, sglang, tensorrt-llm
**Posted:** 2026-08-05

> Owns a multi-cloud GPU infrastructure platform that enables training and inference workloads through self-service orchestration. The role requires deep Kubernetes, distributed systems, GPU networking, and systems programming experience.

## Job Description

## Responsibilities
- Build a self-serve compute platform for launching training jobs and operating inference services without managing GPU provisioning, cluster configuration, or provider-specific infrastructure.
- Own GPU fleet provisioning, lifecycle management, reliability, and capacity integration across cloud providers.
- Build scheduling and placement logic to find available capacity, efficiently pack workloads, and match workloads to appropriate hardware.
- Support long-running distributed training jobs and highly available, low-latency production inference services on the same fleet.
- Write Kubernetes operators and CRDs and manage multi-provider clusters.
- Build fault tolerance, autoscaling, and observability for resilience against node loss, provider failures, and capacity shifts.
- Set technical direction and partner with inference and cloud infrastructure engineers on platform architecture and roadmap.

## Requirements
- Deep Kubernetes experience, including custom operators, CRDs, and multi-cluster federation.
- Experience managing GPU clusters at scale with NVIDIA hardware, CUDA, and high-performance networking such as InfiniBand or RoCE.
- Experience orchestrating compute across multiple clouds, such as CoreWeave, AWS, or GCP.
- Strong distributed-systems fundamentals, including scheduling, resource allocation, and fault tolerance under load.
- Infrastructure and systems-level programming experience in Go, Rust, or C++.
- Experience supporting long-running training jobs and high-availability inference services.
- End-to-end ownership and comfort working independently in ambiguous environments.

## Nice to Have
- Experience with inference-serving stacks such as vLLM, SGLang, or TensorRT-LLM.
- Experience with Slurm or other HPC schedulers.
- GPU kernel experience with CUDA or Triton.
- Production experience with InfiniBand, RoCE, or RDMA.
- Observability experience for ML workloads with Prometheus, Grafana, or Weights & Biases.

## Similar roles

- [Senior Staff Deployment Automation Engineer](https://hotfix.jobs/jobs/e7c5a05a-7667-45ba-8f89-00349f1e9aaa) - Crusoe - San Francisco, CA - $250k – $300k/yr
- [Senior Staff Software Engineer, DC Infrastructure](https://hotfix.jobs/jobs/5e51b5cc-872b-4556-a065-f5f95a363fad) - Crusoe - San Francisco, CA - $250k – $300k/yr
- [Member of Technical Staff](https://hotfix.jobs/jobs/ca5edb5f-58ba-4e0c-a075-25c3fcc61007) - Perplexity - San Francisco, CA - $250k – $405k/yr
- [Staff Engineer, Distributed Storage and HPC & AI Infrastructure](https://hotfix.jobs/jobs/d0fb38e7-169a-4b2d-abb6-6f845d3f381f) - Together AI - San Francisco, CA - $250k – $300k/yr
- [Staff Site Reliability Engineer](https://hotfix.jobs/jobs/e2420bd8-af00-4aa8-9bcb-34da7d97403d) - Zoox - Foster City, CA - $250k – $300k/yr

**Apply:** https://hotfix.jobs/jobs/e2ff6e88-b443-42b7-87ae-33997af4c66f
**Canonical:** https://hotfix.jobs/jobs/e2ff6e88-b443-42b7-87ae-33997af4c66f