Skip to content
FluidstackFluidstackSan Francisco, CA

Infrastructure-as-a-Service Engineer

Build and automate the IaaS layer for large-scale GPU fleets at Fluidstack, including bare-metal provisioning, tenant isolation, lifecycle management, and debugging across compute, network, and storage. Requires production infrastructure automation experience in Go or Python, deep Linux debugging skills, and a software-first approach to operational toil.

204k – 284k/yr
On-site5+ YOEDevOps / SRE

About the role

Role Scope

Build the IaaS layer: provisioning automation, tenant isolation, and lifecycle management for GPU fleets. Automate bare-metal workflows end to end: image, boot, configure, validate, hand over. Debug across compute, network, and storage when tenant workloads misbehave. Write software that removes toil: every manual runbook is a backlog item with your name on it.

What We're Looking For

You've built infrastructure automation running in production, in Go or Python against real systems. You've worked with bare-metal provisioning tooling and know where it lies to you. You debug Linux deeply: kernel, drivers, networking. You treat operational pain as a software bug, and you fix it.

Bonus: GPU systems. Kubernetes internals. IPMI and Redfish. Storage systems.

Skills

GoPythonLinuxKubernetesipmiredfishgpu systemsbare-metal provisioninginfrastructure automation

Similar roles

DevOps / SRE jobs
Okta

Manager, Site Reliability Engineering

OktaSan Francisco, CA

Lead and mentor SRE teams managing Okta's IDaaS platform on AWS, driving DevOps maturity, platform reliability, and tooling across Kubernetes, CI/CD, observability, and automation domains.

204k – 281k/yr
Hybrid3+ YOEDevOps / SRE
Anyscale

Software Engineer, Infrastructure

AnyscaleSan Francisco, CA +1

Software Engineer building scalable control and data plane infrastructure for Anyscale's Ray platform. Design and optimize cluster orchestration, scheduling, Kubernetes deployments, and accelerator support for distributed AI/ML workloads. Requires 3+ years production experience with cloud-native tech, Go/Python, and distributed systems.

202k – 237k/yr
On-site5+ YOEDevOps / SRE
Fluidstack

Software Engineer, Networking

FluidstackSan Francsisco, CA

Build and own end-to-end network fleet health, monitoring, debugging tooling, and automated repair pipelines for one of the world's largest AI datacenter networks. Requires systems thinking, on-call ownership, and fluency with AI coding tools plus Go/Python network automation experience.

208k – 269k/yr
On-site5+ YOEDevOps / SRE
Fluidstack

Software Engineer, Compute

FluidstackSan Francisco, CA +3

Build and own automation, observability, and repair pipelines for one of the world's largest GPU compute fleets. Requires hardware intuition at the firmware/silicon level, on-call ownership, and fluency with AI coding tools to eliminate toil at hyperscale.

208k – 269k/yr
On-site5+ YOEDevOps / SRE
Fluidstack

Site Reliability Engineer, Compute

FluidstackSan Francisco, CA +3

Own end-to-end health, reliability, and automation of a massive GPU compute fleet for AI infrastructure. Build metrics, alerting, repair pipelines, GPU qualification platforms, and low-level BMC/Redfish tooling while driving incidents and using AI coding tools daily.

208k – 269k/yr
On-site5+ YOEDevOps / SRE