Compute Deployment Engineer
Own end-to-end compute deployment and rack qualification for large-scale GPU and accelerator fleets at Fluidstack, from facility handoff through burn-in, validation, and production readiness. Requires deep Linux/out-of-band management experience, hardware automation in Python/Go, data center operations, and methodical failure triage.
About the job
Role Scope
Own compute turn-up from facility availability to ready-for-service: the stretch after the network hands off and before customers run workloads.
Qualify racks at scale: establish firmware baselines, configure BMC and BIOS, run burn-in, and validate at node and cluster level across hundreds of racks per site on GPU and custom accelerator platforms.
Drive qualification through the base-management Kubernetes platform and provisioning stack (discovery, imaging, firmware updates, shared services), burning down qual queues with tooling rather than manual runs.
Triage hardware failures found in qualification: isolate to component, drive RMA and vendor escalation, and feed failure patterns back into the qual gates.
Run turn-up remotely by default, with on-site pulses of roughly a week per data hall as new halls reach facility availability, plus occasional overlapping-site weeks.
Partner with network deployment, ICT, data center operations, and hardware teams during turn-up windows, and support incident response on freshly-live capacity.
Requirements
- Brought up server or GPU fleets at scale, hundreds of nodes or more, and taken them all the way to production.
- Work deep in Linux and out-of-band management: BMC, IPMI, and Redfish.
- Automated hardware workflows in Python or Go rather than clicking through them; turn repeated tasks into software.
- Worked physically in data halls, racking, cabling, and swapping components; effective acting as remote hands or directing them.
- Triage failures methodically across hardware, firmware, and software, isolating the fault to a component.
- Travel for turn-up windows when a new data hall comes online.
Nice-to-Haves
- Kubernetes-based bare-metal provisioning.
- Accelerator platform bringup (NVIDIA, AMD, or custom).
- Burn-in and stress harness design.
- DCIM and inventory tooling.
Skills
Linux, Bmc, Ipmi, Redfish, Python, Go, Kubernetes, Nvidia, Amd
Similar jobs
DevOps / SRE jobsBuild and operate a highly available, multi-region PostgreSQL platform, developing automation, monitoring, disaster recovery, and performance tooling. Requires experience with large-scale PostgreSQL clusters, infrastructure as code, scripting, containers, and observability.
Leads on-site deployment of data center physical infrastructure, managing contractors, performing QA/QC on fiber optics and cabling, and ensuring compliance with standards. Requires 5+ years experience, SME-level fiber optic expertise, bachelor's degree, and 40% travel readiness.
The Python Engineer will improve and operate trading systems, support integrations with asset classes and prime brokers, and handle monitoring, incidents, and performance optimization. The role requires 3+ years of experience, strong Python and Linux skills, and familiarity with market data and order-entry systems.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
The DevOps Engineer will build and operate reliable infrastructure, deployment workflows, and observability for data pipelines and AI/ML systems. The role requires at least three years of DevOps, SRE, or infrastructure experience plus strong cloud, Terraform, containerization, and MLOps expertise.