# Staff Software Engineer, Inference / Compute Infrastructure Engineering

**Company:** [Together AI](https://hotfix.jobs/companies/together-ai)
**Location:** Remote
**Role:** DevOps / SRE
**Experience:** 7+ years
**Skills:** Go, Python, Rust, Kubernetes, Kubernetes Controllers, Crds, Temporal, Cadence, Event-Driven Systems, Kafka, Nats, SQS, CUDA, Nccl, InfiniBand
**Posted:** 2026-08-30

> Build and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.

## Job Description

## Responsibilities
- Build a provisioning state machine modeling the lifecycle of physical hosts, from discovery and inference bring-up through GPU driver/CUDA installation, health validation, and decommissioning or RMA, using explicit, versioned states and transitions.
- Design declarative self-service APIs and a control plane for requesting, scaling, and tearing down inference clusters.
- Automate self-healing by detecting degraded or failed nodes, safely draining them, triggering repair or replacement, and returning healthy capacity to the pool.
- Own pipeline reliability through idempotency, retries, rollback, and drift detection.
- Partner with inference and ML platform teams to encode cluster topology, interconnect, and scheduling constraints as platform abstractions.
- Apply software engineering practices including strong typing, automated testing, code review, versioning, and CI/CD for infrastructure code.

## Requirements
- Strong software engineering background in Go, Python, Rust, or a similar language.
- Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent.
- Experience building control planes or orchestration systems that model state and reconcile it over time, such as Kubernetes controllers, operators, custom reconciliation loops, or workflow engines.
- Experience designing event-driven systems using message queues, event streams, or pub/sub rather than polling or cron-driven scripts.
- Experience building internal platforms or APIs used by other engineering teams, with attention to developer experience.

## Nice to Have
- Exposure to bare-metal provisioning, including PXE/iPXE, Redfish/IPMI, or BMC.
- Networking fundamentals such as VLANs, BGP, or fabric design.
- Experience with GPU or accelerator infrastructure.
- Familiarity with GPU cluster stacks such as NCCL, CUDA, and InfiniBand/RoCE.
- Experience at a hyperscaler, GPU cloud, or datacenter-scale infrastructure organization.
- Systems programming in Rust or Go.

## Similar jobs

- [Staff Software Engineer, Inference / Compute Infrastructure Engineering](https://hotfix.jobs/jobs/25af7190-dc4d-4d85-b2b7-5b41d1dc4405) - Together AI - London, United Kingdom
- [Staff Software Engineer, Observability & Profiling](https://hotfix.jobs/jobs/be415182-687b-4ae9-a1c2-4b70b877514b) - Anthropic - London, United Kingdom - £325k – £390k/yr
- [Staff Software Engineer: Platform SRE](https://hotfix.jobs/jobs/d1231c10-b3f1-48c5-a137-495d17d4f69a) - Flexport - Amsterdam, Netherlands
- [Staff Software Engineer, AI Reliability Engineering](https://hotfix.jobs/jobs/66b246fd-6b0e-4283-b1ff-c460780b61a2) - Anthropic - London, United Kingdom - £325k – £390k/yr
- [Staff Engineer, Digital Infrastructure](https://hotfix.jobs/jobs/ad842bf1-8db2-4951-b37a-ea4e7c1f4977) - Shield AI - London, United Kingdom

**Apply:** https://hotfix.jobs/jobs/b1b1bdb1-d3f0-47cd-b94b-7781be7399fc
**Canonical:** https://hotfix.jobs/jobs/b1b1bdb1-d3f0-47cd-b94b-7781be7399fc