Skip to content

Staff Software Engineer, Inference / Compute Infrastructure Engineering

Build and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.

About the job

Responsibilities

  • Build a provisioning state machine modeling the full lifecycle of physical hosts, from discovery and inference bring-up through GPU driver/CUDA installation, health validation, and decommissioning/RMA, using explicit, versioned states and transitions.
  • Design and implement declarative self-service APIs and a Kubernetes-native control plane for requesting, scaling, and tearing down inference clusters.
  • Automate self-healing by detecting degraded or failed nodes, safely draining them, triggering repair or replacement, and automatically restoring healthy capacity.
  • Ensure pipeline reliability through idempotency, retries, rollback, and drift detection.
  • Partner with the inference and ML platform teams to encode cluster topology, interconnect, and scheduling requirements as platform abstractions.
  • Develop controllers, reconciliation loops, APIs/CLIs, event-driven health and remediation systems, and GPU utilization optimizations including defragmentation, rebalancing, scheduling, bin-packing, and right-sizing.
  • Operate and support the software in production, with strong typing, automated tests, code review, versioning, and CI/CD for infrastructure code.

Requirements

  • Strong software engineering experience in Go, Python, Rust, or a similar language.
  • Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent.
  • Experience building control planes or orchestration systems that model state and reconcile it over time, such as Kubernetes controllers/operators, custom reconciliation loops, or workflow engines.
  • Experience designing event-driven systems using message queues, event streams, or pub/sub.
  • Experience building internal platforms or APIs used by other engineering teams, with attention to developer experience.

Nice to Have

  • Exposure to bare-metal provisioning, including PXE/iPXE, Redfish/IPMI, or BMC.
  • Networking fundamentals such as VLANs, BGP, or fabric design.
  • GPU or accelerator infrastructure experience.
  • Experience with NCCL, CUDA, InfiniBand, or RoCE.
  • Experience at a hyperscaler, GPU cloud, or datacenter-scale infrastructure organization.
  • Systems programming in Rust or Go.

Skills

Kubernetes, Go, Python, Rust, Temporal, Cadence, Kafka, Nats, Amazon Sqs, CUDA, Nccl, InfiniBand, Roce, Pxe/Ipxe, Redfish/Ipmi

Anthropic

Anthropic

London, United Kingdom

Staff Software Engineer, Observability & Profiling
£325k+/yrHybrid10+ YOEDevOps / SRE

Build and operate foundational observability infrastructure spanning telemetry pipelines, profiling, tracing, and diagnostic tooling across large-scale compute clusters. The role requires deep systems-level experience and 10+ years of relevant industry experience.

Flexport

Flexport

Amsterdam, Netherlands

Staff Software Engineer: Platform SRE
No salary listedHybrid10+ YOEDevOps / SRE

Leads the design, operation, and evolution of Flexport’s cloud infrastructure, platform tooling, observability, and incident response systems. Requires 10+ years of software, SRE, or infrastructure engineering experience, deep AWS expertise, and strong Terraform and automation skills.

Together AI

Together AI

London, United Kingdom
Staff Software Engineer, Inference / Compute Infrastructure Engineering
No salary listedRemote7+ YOEDevOps / SRE

Build and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.

Anthropic

Anthropic

London, United Kingdom

Staff Software Engineer, AI Reliability Engineering
£325k+/yrHybrid7+ YOEDevOps / SRE

Leads reliability engineering for critical AI serving systems, spanning SLOs, observability, high availability, and incident response. Requires strong distributed-systems or infrastructure experience, with model-serving, accelerator, networking, and resilience-testing expertise valued.

Shield AI

Shield AI

London, United Kingdom

Staff Engineer, Digital Infrastructure
No salary listedOn-site7+ YOEDevOps / SRE

Staff Engineer responsible for deploying, integrating, maintaining, and developing an AI training factory across isolated environments. The role requires 7+ years of related experience, cloud and Kubernetes expertise, Linux networking knowledge, application support skills, and automation experience.