Skip to content

Staff Software Engineer, GPU Infrastructure Lifecycle Management

Build and own software state machines and control planes that automate the full lifecycle of GPU infrastructure from bare metal provisioning to running AI inference clusters. Requires strong software engineering experience with orchestration, reconciliation loops, and event-driven systems.

About the job

Responsibilities

  • Build the provisioning state machine: design and implement the software that models the full lifecycle of a physical host from discovery, inference bring-up to GPU driver/CUDA stack, health validation, and decommission/RMA — as explicit, versioned states and transitions.
  • Build the self-service API: design declarative APIs and a control plane so the inference team can request, scale, and tear down inference clusters with one API call — no ticket, no human in the loop.
  • Automate self-healing: detect degraded or failed nodes, drain them safely, trigger repair or replacement, and reintroduce healthy capacity into the pool automatically.
  • Own reliability of the pipeline: idempotency, retries, rollback, and drift detection so the provisioning system is as dependable as any other production service.
  • Partner with the inference/ML platform team: understand the cluster shapes they need — topology, interconnect, scheduling constraints — and encode them as first-class abstractions in the platform.
  • Engineer it like software: strong typing, automated tests, code review, versioning, and CI/CD for infrastructure code — this is a product, not a collection of Ansible playbooks.

Requirements

  • Strong software engineering background in Go, Python, Rust, or similar.
  • Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent to run long-lived, manifest-driven workflows that survive failures and resume mid-execution.
  • Experience building software control planes or orchestration systems that model state and reconcile it over time (e.g., Kubernetes controllers/operators, custom reconciliation loops, workflow engines).
  • Experience with event-driven systems — designing and building software around message queues, event streams, or pub/sub (e.g., Kafka, NATS, SQS).
  • A product mindset. You’ve built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship.

Nice-to-Haves

  • Exposure to bare-metal provisioning (PXE/iPXE, Redfish/IPMI, BMC) and/or networking fundamentals (VLANs, BGP, fabric design), or GPU/accelerator infrastructure.
  • Experience with GPU cluster software stacks (NCCL, CUDA, InfiniBand/RoCE).
  • Prior work at a hyperscaler, GPU cloud, or datacenter-scale infrastructure organization.
  • Systems programming in Rust or Go.

Compensation

  • US base salary range: $240,000 - $280,000 + equity + benefits.

Skills

Go, Python, Rust, Temporal, Cadence, Kubernetes, Kafka, Nats, SQS, CUDA, Nccl, InfiniBand, Redfish, Ipmi

Skydio

Skydio

San Mateo, CA
Staff Site Reliability Engineer
$240k+/yrRemote8+ YOEDevOps / SRE

Owns and scales production cloud infrastructure across Kubernetes/EKS, AWS, Terraform, CI/CD, networking, and observability. The role requires 8+ years of infrastructure experience, strong Kubernetes operations expertise, and depth in reliability or scaling challenges.

Datadog

Datadog

Boston, MA
Staff Engineer - Cloud Networks
$244k+/yrHybrid7+ YOEDevOps / SRE

Leads the technical direction, design, and operation of large-scale multi-cloud network infrastructure, with a focus on connectivity, reliability, performance, and cost efficiency. Requires deep BGP and software-defined networking expertise plus strong software development and production operations experience.

Polymarket

Polymarket

New York, NY

Staff Infrastructure Engineer
$250k+/yrOn-site7+ YOEDevOps / SRE

Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.

Crusoe

Crusoe

San Francisco, CA
Senior Staff Deployment Automation Engineer
$250k+/yrOn-site12+ YOEDevOps / SRE

Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.

Crusoe

Crusoe

San Francisco, CA
Senior Staff Software Engineer, DC Infrastructure
$250k+/yrOn-site7+ YOEDevOps / SRE

Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.