Skip to content
FluidstackFluidstackSan Francisco, CA

Compute Deployment Engineer

Own end-to-end compute deployment and rack qualification for large-scale GPU and accelerator fleets at Fluidstack, from facility handoff through burn-in, validation, and production readiness. Requires deep Linux/out-of-band management experience, hardware automation in Python/Go, data center operations, and methodical failure triage.

150k – 250k/yr
Hybrid5+ YOEDevOps / SRE

About the role

Role Scope

Own compute turn-up from facility availability to ready-for-service: the stretch after the network hands off and before customers run workloads.

Qualify racks at scale: establish firmware baselines, configure BMC and BIOS, run burn-in, and validate at node and cluster level across hundreds of racks per site on GPU and custom accelerator platforms.

Drive qualification through the base-management Kubernetes platform and provisioning stack (discovery, imaging, firmware updates, shared services), burning down qual queues with tooling rather than manual runs.

Triage hardware failures found in qualification: isolate to component, drive RMA and vendor escalation, and feed failure patterns back into the qual gates.

Run turn-up remotely by default, with on-site pulses of roughly a week per data hall as new halls reach facility availability, plus occasional overlapping-site weeks.

Partner with network deployment, ICT, data center operations, and hardware teams during turn-up windows, and support incident response on freshly-live capacity.

Requirements

  • Brought up server or GPU fleets at scale, hundreds of nodes or more, and taken them all the way to production.
  • Work deep in Linux and out-of-band management: BMC, IPMI, and Redfish.
  • Automated hardware workflows in Python or Go rather than clicking through them; turn repeated tasks into software.
  • Worked physically in data halls, racking, cabling, and swapping components; effective acting as remote hands or directing them.
  • Triage failures methodically across hardware, firmware, and software, isolating the fault to a component.
  • Travel for turn-up windows when a new data hall comes online.

Nice-to-Haves

  • Kubernetes-based bare-metal provisioning.
  • Accelerator platform bringup (NVIDIA, AMD, or custom).
  • Burn-in and stress harness design.
  • DCIM and inventory tooling.

Skills

LinuxbmcipmiredfishPythonGoKubernetesnvidiaamd

Similar roles

DevOps / SRE jobs
Encord

DevOps Engineer

EncordSan Francisco, CA

DevOps Engineer embedded in platform teams to build and operate scalable AI infrastructure on GCP/AWS. Own CI/CD, Kubernetes, observability, reliability (SLIs/SLOs), automation, and performance at petabyte scale. Requires 4-5 years production DevOps/SRE experience.

150k – 170k/yr
On-site4+ YOEDevOps / SRE
Fluidstack

Network Engineer, BMS/EPMS Networks

FluidstackNew York, NY

Own and design OT facility networks for BMS, EPMS, and controls traffic in AI data centers. Requires experience building industrial/OT networks, deep knowledge of controls protocols like BACnet/IP, and implementing segmentation that works with operations teams.

150k – 203k/yr
On-site5+ YOEDevOps / SRE
Runpod

Site Reliability Engineer

RunpodUnited States

Site Reliability Engineer responsible for defining SLIs/SLOs, leading incident response, building observability (Prometheus/Grafana), automating toil, and driving production readiness for Runpod's AI cloud platform. Requires 5+ years SRE experience, strong Linux/distributed systems knowledge, and scripting skills.

150k – 200k/yr
Remote5+ YOEDevOps / SRE
Turion Space

DevSecOps Engineer

Turion SpaceIrvine, CA

DevSecOps Engineer designing, developing, and maintaining reliable C++/Python software systems with a focus on CI/CD automation, build systems, and developer tooling for aerospace and defense applications. Requires 3+ years experience, proficiency in modern C++, CMake, and CI/CD pipelines.

150k – 213k/yr
Hybrid3+ YOEDevOps / SRE
Clear Street

Software Engineer

Clear StreetNew York, NY

Platform Engineer building self-service internal developer platforms, reusable services, intelligent CI/CD, and automation to improve developer productivity and software delivery at a capital markets fintech. Requires 5+ years experience, strong software engineering skills in Go/Python/Java, deep cloud-native and Kubernetes expertise, and a product mindset for internal tools.

150k – 200k/yr
Hybrid5+ YOEDevOps / SRE