Skip to content
FluidstackFluidstackSan Francisco, CA

Software Engineer, Compute

Build and own automation, observability, and repair pipelines for one of the world's largest GPU compute fleets. Requires hardware intuition at the firmware/silicon level, on-call ownership, and fluency with AI coding tools to eliminate toil at hyperscale.

208k – 269k/yr
On-site5+ YOEDevOps / SRE

About the role

Responsibilities

  • Own compute fleet health end to end. Build the metrics pipelines, alerting, and unified health view that tell you the true state of every GPU in production — across Kubernetes-orchestrated workloads and bare metal, at scale.
  • Turn deployment/repair into a pipeline, not a procedure. Build and own the automation that takes a compute failure from detection through triage, parts management, and return to service. No one-off scripts, no heroics.
  • Design and expand the GPU qualification platform. Burn-in, performance baselining, and NPI execution for every new GPU generation. Define what "good" looks like before hardware goes into production.
  • Own Redfish and BMC tooling. Firmware-level telemetry, log collection at fleet scale, and the low-level access layer that repair automation and health tooling depend on.
  • Own end-to-end reliability, scalability, and operation of the compute fleet at-scale. Build aggressive automation, tooling, and incident discipline for one of the largest GPU fleets in the world.

Requirements

  • Treat toil as a bug. Manual steps in a repair workflow are a backlog item, not a job description.
  • Instinct for hardware. Comfortable reasoning about failure modes at the firmware and silicon level, not just the software stack above it.
  • Move toward ambiguity, not away from it. Walk into the fog, build the map, and explain it to everyone else.
  • Learn at a steep slope. Reach real competence in an unfamiliar domain fast.
  • Carry a pager without flinching. Run the incident, write the postmortem, fix the systemic cause, and move on.
  • Fluent with AI tooling. LLM APIs, MCP servers, and agentic frameworks; drive Claude Code, Cursor, or similar every day.
  • Shipped production automation that other teams depend on, and comfortable in any language using AI coding tools.

Nice-to-Haves

  • Hardware lifecycle management and RMA automation.
  • BMC/Redfish or IPMI tooling.
  • GPU qualification or burn-in frameworks.
  • Workflow and orchestration engines (Temporal, Cadence).
  • Metrics and alerting pipelines (Prometheus, Grafana).
  • Go or Python.

Skills

KubernetesredfishbmcPrometheusGrafanaGoPythonTemporalcadenceLLM APIsgpu qualification

Similar roles

DevOps / SRE jobs
Fluidstack

Software Engineer, Networking

FluidstackSan Francsisco, CA

Build and own end-to-end network fleet health, monitoring, debugging tooling, and automated repair pipelines for one of the world's largest AI datacenter networks. Requires systems thinking, on-call ownership, and fluency with AI coding tools plus Go/Python network automation experience.

208k – 269k/yr
On-site5+ YOEDevOps / SRE
Fluidstack

Site Reliability Engineer, Compute

FluidstackSan Francisco, CA +3

Own end-to-end health, reliability, and automation of a massive GPU compute fleet for AI infrastructure. Build metrics, alerting, repair pipelines, GPU qualification platforms, and low-level BMC/Redfish tooling while driving incidents and using AI coding tools daily.

208k – 269k/yr
On-site5+ YOEDevOps / SRE
Fluidstack

Network Software Engineer

FluidstackSan Francisco, CA +3

Build and own end-to-end network monitoring, debugging tooling, repair automation pipelines, and qualification frameworks for Fluidstack's massive AI datacenter fleet. Requires systems thinking, on-call ownership, Go/Python automation experience, network protocols expertise, and daily use of AI coding tools.

208k – 269k/yr
On-site5+ YOEDevOps / SRE
Fluidstack

Network Automation Engineer

FluidstackSan Francisco, CA +3

Own end-to-end network fleet health, monitoring, debugging tooling, and automated repair pipelines for massive AI datacenter infrastructure. Requires systems thinking, building automation in Go/Python, experience with optical networks and protocols like gNMI/gRPC/NETCONF/SONiC, and daily use of AI coding tools.

208k – 269k/yr
On-site5+ YOEDevOps / SRE
Fluidstack

Distributed Systems Engineer

FluidstackSan Francisco, CA +3

Build and own the observability platform, production control plane, and fleet state as source of truth for a hyperscale GPU fleet powering AI infrastructure. Requires shipping scalable production services, on-call ownership, and comfort with AI coding tools; distributed systems and observability experience preferred.

208k – 269k/yr
On-site5+ YOEDevOps / SRE