Staff Software Engineer, Inference / Compute Infrastructure Engineering
Build and operate declarative control planes, durable workflows, and self-healing systems that provision and manage GPU inference infrastructure. The role requires strong software engineering, reconciliation or orchestration experience, and event-driven systems expertise.
About the job
Responsibilities
- Build a provisioning state machine for the full physical-host lifecycle, from discovery and inference bring-up through GPU driver/CUDA installation, health validation, decommissioning, and RMA.
- Design declarative self-service APIs and a control plane for requesting, scaling, and tearing down inference clusters.
- Automate self-healing by detecting degraded or failed nodes, safely draining them, triggering repair or replacement, and restoring healthy capacity.
- Ensure pipeline reliability through idempotency, retries, rollback, and drift detection.
- Partner with inference and ML platform teams to model cluster topology, interconnects, and scheduling constraints.
- Apply production software practices, including strong typing, automated tests, code review, versioning, and CI/CD.
Requirements
- Strong software engineering experience with Go, Python, Rust, or similar languages.
- Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent.
- Experience building control planes or orchestration systems that model state and reconcile it over time, such as Kubernetes controllers, operators, or custom reconciliation loops.
- Experience designing event-driven systems around message queues, event streams, or pub/sub.
- Experience building internal platforms or APIs for engineering teams, with attention to developer experience.
Nice to have
- Experience with bare-metal provisioning, including PXE/iPXE, Redfish/IPMI, or BMC.
- Networking fundamentals, including VLANs, BGP, or fabric design.
- GPU or accelerator infrastructure experience.
- Experience with NCCL, CUDA, InfiniBand, or RoCE.
- Experience at a hyperscaler, GPU cloud, or datacenter-scale infrastructure organization.
- Systems programming in Rust or Go.
Skills
Go, Python, Rust, Temporal, Cadence, Kubernetes, Event-Driven Systems, Kafka, Nats, Amazon Sqs, Pxe, Ipxe, Redfish, CUDA
Similar jobs
DevOps / SRE jobsBuild and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.
Build and operate foundational observability infrastructure spanning telemetry pipelines, profiling, tracing, and diagnostic tooling across large-scale compute clusters. The role requires deep systems-level experience and 10+ years of relevant industry experience.
Leads the design, operation, and evolution of Flexport’s cloud infrastructure, platform tooling, observability, and incident response systems. Requires 10+ years of software, SRE, or infrastructure engineering experience, deep AWS expertise, and strong Terraform and automation skills.
Build and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.
Leads reliability engineering for critical AI serving systems, spanning SLOs, observability, high availability, and incident response. Requires strong distributed-systems or infrastructure experience, with model-serving, accelerator, networking, and resilience-testing expertise valued.