# Software Engineer, Compute

**Company:** [Fluidstack](https://hotfix.jobs/companies/fluidstack)
**Location:** San Francisco, CA, New York, NY, Austin, TX, Seattle, WA
**Role:** DevOps / SRE
**Salary:** $208k – $269k/yr
**Experience:** 5+ years
**Skills:** Kubernetes, redfish, bmc, Prometheus, Grafana, Go, Python, Temporal, cadence, LLM APIs, gpu qualification
**Posted:** 2026-07-20

> Build and own automation, observability, and repair pipelines for one of the world's largest GPU compute fleets. Requires hardware intuition at the firmware/silicon level, on-call ownership, and fluency with AI coding tools to eliminate toil at hyperscale.

## Job Description

## Responsibilities
- Own compute fleet health end to end. Build the metrics pipelines, alerting, and unified health view that tell you the true state of every GPU in production — across Kubernetes-orchestrated workloads and bare metal, at scale.
- Turn deployment/repair into a pipeline, not a procedure. Build and own the automation that takes a compute failure from detection through triage, parts management, and return to service. No one-off scripts, no heroics.
- Design and expand the GPU qualification platform. Burn-in, performance baselining, and NPI execution for every new GPU generation. Define what "good" looks like before hardware goes into production.
- Own Redfish and BMC tooling. Firmware-level telemetry, log collection at fleet scale, and the low-level access layer that repair automation and health tooling depend on.
- Own end-to-end reliability, scalability, and operation of the compute fleet at-scale. Build aggressive automation, tooling, and incident discipline for one of the largest GPU fleets in the world.

## Requirements
- Treat toil as a bug. Manual steps in a repair workflow are a backlog item, not a job description.
- Instinct for hardware. Comfortable reasoning about failure modes at the firmware and silicon level, not just the software stack above it.
- Move toward ambiguity, not away from it. Walk into the fog, build the map, and explain it to everyone else.
- Learn at a steep slope. Reach real competence in an unfamiliar domain fast.
- Carry a pager without flinching. Run the incident, write the postmortem, fix the systemic cause, and move on.
- Fluent with AI tooling. LLM APIs, MCP servers, and agentic frameworks; drive Claude Code, Cursor, or similar every day.
- Shipped production automation that other teams depend on, and comfortable in any language using AI coding tools.

## Nice-to-Haves
- Hardware lifecycle management and RMA automation.
- BMC/Redfish or IPMI tooling.
- GPU qualification or burn-in frameworks.
- Workflow and orchestration engines (Temporal, Cadence).
- Metrics and alerting pipelines (Prometheus, Grafana).
- Go or Python.

## Similar roles

- [Software Engineer, Networking](https://hotfix.jobs/jobs/c083d3e7-a957-45f7-9406-c00edae189b3) - Fluidstack - San Francsisco, CA - $208k – $269k/yr
- [Site Reliability Engineer, Compute](https://hotfix.jobs/jobs/4374777f-3589-4aac-bd7c-61c791f5a467) - Fluidstack - San Francisco, CA - $208k – $269k/yr
- [Network Software Engineer](https://hotfix.jobs/jobs/3691ff92-fa54-43af-8c26-577a95debbae) - Fluidstack - San Francisco, CA - $208k – $269k/yr
- [Network Automation Engineer](https://hotfix.jobs/jobs/98f17fdc-79fd-4af9-a06f-5d58edb06019) - Fluidstack - San Francisco, CA - $208k – $269k/yr
- [Distributed Systems Engineer](https://hotfix.jobs/jobs/8505fdd9-bfbf-438a-9c6f-64b2d80f7c5c) - Fluidstack - San Francisco, CA - $208k – $269k/yr

**Apply:** https://hotfix.jobs/jobs/a47fd61c-f69f-4b10-ad03-bd207719bb21
**Canonical:** https://hotfix.jobs/jobs/a47fd61c-f69f-4b10-ad03-bd207719bb21