# Compute Deployment Engineer

**Company:** [Fluidstack](https://hotfix.jobs/companies/fluidstack)
**Location:** San Francisco, CA, New York, NY
**Role:** DevOps / SRE
**Salary:** $150k – $250k/yr
**Experience:** 5+ years
**Skills:** Linux, bmc, ipmi, redfish, Python, Go, Kubernetes, nvidia, amd
**Posted:** 2026-07-17

> Own end-to-end compute deployment and rack qualification for large-scale GPU and accelerator fleets at Fluidstack, from facility handoff through burn-in, validation, and production readiness. Requires deep Linux/out-of-band management experience, hardware automation in Python/Go, data center operations, and methodical failure triage.

## Job Description

## Role Scope
Own compute turn-up from facility availability to ready-for-service: the stretch after the network hands off and before customers run workloads.

Qualify racks at scale: establish firmware baselines, configure BMC and BIOS, run burn-in, and validate at node and cluster level across hundreds of racks per site on GPU and custom accelerator platforms.

Drive qualification through the base-management Kubernetes platform and provisioning stack (discovery, imaging, firmware updates, shared services), burning down qual queues with tooling rather than manual runs.

Triage hardware failures found in qualification: isolate to component, drive RMA and vendor escalation, and feed failure patterns back into the qual gates.

Run turn-up remotely by default, with on-site pulses of roughly a week per data hall as new halls reach facility availability, plus occasional overlapping-site weeks.

Partner with network deployment, ICT, data center operations, and hardware teams during turn-up windows, and support incident response on freshly-live capacity.

## Requirements
- Brought up server or GPU fleets at scale, hundreds of nodes or more, and taken them all the way to production.
- Work deep in Linux and out-of-band management: BMC, IPMI, and Redfish.
- Automated hardware workflows in Python or Go rather than clicking through them; turn repeated tasks into software.
- Worked physically in data halls, racking, cabling, and swapping components; effective acting as remote hands or directing them.
- Triage failures methodically across hardware, firmware, and software, isolating the fault to a component.
- Travel for turn-up windows when a new data hall comes online.

## Nice-to-Haves
- Kubernetes-based bare-metal provisioning.
- Accelerator platform bringup (NVIDIA, AMD, or custom).
- Burn-in and stress harness design.
- DCIM and inventory tooling.

## Similar roles

- [DevOps Engineer](https://hotfix.jobs/jobs/d6ef1d0e-b1aa-4f8e-a5a0-ecc6a68be5ed) - Encord - San Francisco, CA - $150k – $170k/yr
- [Network Engineer, BMS/EPMS Networks](https://hotfix.jobs/jobs/ff035650-d0e1-42f0-98ce-40cb5d07708f) - Fluidstack - New York, NY - $150k – $203k/yr
- [Site Reliability Engineer](https://hotfix.jobs/jobs/31ec0a81-7790-4ee1-94de-5729935ed7e5) - Runpod - Remote - $150k – $200k/yr
- [DevSecOps Engineer](https://hotfix.jobs/jobs/67398973-784d-4b66-ab52-901744181150) - Turion Space - Irvine, CA - $150k – $213k/yr
- [Software Engineer](https://hotfix.jobs/jobs/ba2266bc-5c0a-4708-b7d4-e006c9fec0cd) - Clear Street - New York, NY - $150k – $200k/yr

**Apply:** https://hotfix.jobs/jobs/344acfc4-263d-42c3-95f2-14f7bd1a4511
**Canonical:** https://hotfix.jobs/jobs/344acfc4-263d-42c3-95f2-14f7bd1a4511