Staff Software Engineer, Inference / Compute Infrastructure Engineering
Build and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.
About the job
Responsibilities
- Build a provisioning state machine modeling the lifecycle of physical hosts, from discovery and inference bring-up through GPU driver/CUDA installation, health validation, and decommissioning or RMA, using explicit, versioned states and transitions.
- Design declarative self-service APIs and a control plane for requesting, scaling, and tearing down inference clusters.
- Automate self-healing by detecting degraded or failed nodes, safely draining them, triggering repair or replacement, and returning healthy capacity to the pool.
- Own pipeline reliability through idempotency, retries, rollback, and drift detection.
- Partner with inference and ML platform teams to encode cluster topology, interconnect, and scheduling constraints as platform abstractions.
- Apply software engineering practices including strong typing, automated testing, code review, versioning, and CI/CD for infrastructure code.
Requirements
- Strong software engineering background in Go, Python, Rust, or a similar language.
- Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent.
- Experience building control planes or orchestration systems that model state and reconcile it over time, such as Kubernetes controllers, operators, custom reconciliation loops, or workflow engines.
- Experience designing event-driven systems using message queues, event streams, or pub/sub rather than polling or cron-driven scripts.
- Experience building internal platforms or APIs used by other engineering teams, with attention to developer experience.
Nice to Have
- Exposure to bare-metal provisioning, including PXE/iPXE, Redfish/IPMI, or BMC.
- Networking fundamentals such as VLANs, BGP, or fabric design.
- Experience with GPU or accelerator infrastructure.
- Familiarity with GPU cluster stacks such as NCCL, CUDA, and InfiniBand/RoCE.
- Experience at a hyperscaler, GPU cloud, or datacenter-scale infrastructure organization.
- Systems programming in Rust or Go.
Skills
Go, Python, Rust, Kubernetes, Kubernetes Controllers, Crds, Temporal, Cadence, Event-Driven Systems, Kafka, Nats, SQS, CUDA, Nccl, InfiniBand
Similar jobs
DevOps / SRE jobsBuild and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.
Build and operate foundational observability infrastructure spanning telemetry pipelines, profiling, tracing, and diagnostic tooling across large-scale compute clusters. The role requires deep systems-level experience and 10+ years of relevant industry experience.
Leads the design, operation, and evolution of Flexport’s cloud infrastructure, platform tooling, observability, and incident response systems. Requires 10+ years of software, SRE, or infrastructure engineering experience, deep AWS expertise, and strong Terraform and automation skills.
Leads reliability engineering for critical AI serving systems, spanning SLOs, observability, high availability, and incident response. Requires strong distributed-systems or infrastructure experience, with model-serving, accelerator, networking, and resilience-testing expertise valued.
Staff Engineer responsible for deploying, integrating, maintaining, and developing an AI training factory across isolated environments. The role requires 7+ years of related experience, cloud and Kubernetes expertise, Linux networking knowledge, application support skills, and automation experience.