# Software Engineer, GPU Infrastructure

**Company:** [Fluidstack](https://hotfix.jobs/companies/fluidstack)
**Location:** San Francsisco, CA, New York, NY, Austin, TX, Seattle, WA
**Role:** DevOps / SRE
**Salary:** $175k – $300k/yr
**Experience:** 5+ years
**Skills:** Kubernetes, redfish, bmc, Prometheus, Grafana, Go, Python, Temporal, cadence, LLM APIs, gpu infrastructure
**Posted:** 2026-07-20

> Build and own automation, observability, and repair pipelines for one of the world's largest GPU compute fleets at hyperscale. Requires strong production engineering experience, hardware intuition at the firmware/silicon level, on-call ownership, and fluency with AI coding tools.

## Job Description

## Production Engineering Team

Examples of key exciting problems the team is working on:
- Build the repair pipeline that keeps pace with a fleet of 10s to 100s of GWs: at our scale, a GPU failure isn't a ticket. It's a throughput problem. We're building the automation that takes a chip from fault detection through triage, RMA, and return to service without human intervention.
- Qualify every new GPU generation inside a 6-month build window: our platform covers burn-in, performance baselining, and NPI execution. It has to define "production-ready" before a site goes live, not after. New hardware gets certified at speeds unheard of in the industry.
- Migrate live compute at construction speed: we're converting clusters across production sites simultaneously, bringing new sites online, and making Kubernetes-orchestrated bare metal sustainable at the pace we're building – multiple GW annually.
- See and own the entire fleet in real time, at any scale: build the observability and orchestration layer that makes hyperscale AI compute actually operable. Debug, tune, and performance-test infrastructure that grows by another site every few months.

## Role Scope
- Own compute fleet health end to end. Build the metrics pipelines, alerting, and unified health view that tell you the true state of every GPU in production — across Kubernetes-orchestrated workloads and bare metal, at scale.
- Turn deployment/repair into a pipeline, not a procedure. Build and own the automation that takes a compute failure from detection through triage, parts management, and return to service. No one-off scripts, no heroics.
- Design and expand the GPU qualification platform. Burn-in, performance baselining, and NPI execution for every new GPU generation. You define what "good" looks like before hardware goes into production.
- Own Redfish and BMC tooling. Firmware-level telemetry, log collection at fleet scale, and the low-level access layer that repair automation and health tooling depend on.
- Own end-to-end reliability, scalability, and operation of the compute fleet at-scale. Fluidstack is building one of the largest GPU fleets in the world and that can only be accomplished with aggressive automation, tooling, and incident discipline.

## What We're Looking For
- You treat toil as a bug. Manual steps in a repair workflow are a backlog item, not a job description.
- You have an instinct for hardware. You're comfortable reasoning about failure modes at the firmware and silicon level, not just the software stack above it.
- You move toward ambiguity, not away from it. You walk into the fog, build the map, and explain it to everyone else.
- You learn at a steep slope. You reach real competence in an unfamiliar domain fast. We value this over existing expertise.
- You carry a pager without flinching. You run the incident, write the postmortem, fix the systemic cause, and move on.
- You're fluent with AI tooling. LLM APIs, MCP servers, and agentic frameworks, and you drive Claude Code, Cursor, or similar every day.
- You've shipped production automation that other teams depend on, and you're comfortable in any language using AI coding tools.

**Bonus:**
- Hardware lifecycle management and RMA automation.
- BMC/Redfish or IPMI tooling.
- GPU qualification or burn-in frameworks.
- Workflow and orchestration engines (Temporal, Cadence).
- Metrics and alerting pipelines (Prometheus, Grafana).
- Go or Python.

## Salary & Benefits
- Competitive total compensation package (salary + equity).
- Retirement or pension plan, in line with local norms.
- Health, dental, and vision insurance.
- Generous PTO policy, in line with local norms.

The base salary range for this position is $175,000 - $300,000 per year, depending on experience, skills, qualifications, and location.

## Similar roles

- [Software Engineer, Cloud Infrastructure](https://hotfix.jobs/jobs/936403e0-45ba-4cdc-ad35-13da95bce47b) - Fluidstack - San Francsisco, CA - $175k – $300k/yr
- [Production Engineer, Network](https://hotfix.jobs/jobs/4ed1b583-f227-4746-958f-535c2b4acf48) - Fluidstack - Austin, TX - $175k – $300k/yr
- [Production Engineer, Compute](https://hotfix.jobs/jobs/fd05f1e5-25dc-4cc1-838c-121d9ad18c39) - Fluidstack - San Francisco, CA - $175k – $300k/yr
- [SWE - Backend Infrastructure Engineer](https://hotfix.jobs/jobs/2fb5751d-dec2-4824-81fb-e88e0bebf1e6) - Sesame - San Francisco, CA - $175k – $280k/yr
- [Site Reliability Engineer](https://hotfix.jobs/jobs/d04240e3-4b33-406e-acc5-592de1a1fa94) - Blaxel - San Francisco, CA - $175k – $250k/yr

**Apply:** https://hotfix.jobs/jobs/f636fc6b-b6a5-4eb2-86ac-270e27a733c3
**Canonical:** https://hotfix.jobs/jobs/f636fc6b-b6a5-4eb2-86ac-270e27a733c3