Role Scope
Own compute turn-up from facility availability to ready-for-service: the stretch after the network hands off and before customers run workloads.
Qualify racks at scale: establish firmware baselines, configure BMC and BIOS, run burn-in, and validate at node and cluster level across hundreds of racks per site on GPU and custom accelerator platforms.
Drive qualification through the base-management Kubernetes platform and provisioning stack (discovery, imaging, firmware updates, shared services), burning down qual queues with tooling rather than manual runs.
Triage hardware failures found in qualification: isolate to component, drive RMA and vendor escalation, and feed failure patterns back into the qual gates.
Run turn-up remotely by default, with on-site pulses of roughly a week per data hall as new halls reach facility availability, plus occasional overlapping-site weeks.
Partner with network deployment, ICT, data center operations, and hardware teams during turn-up windows, and support incident response on freshly-live capacity.
Requirements
- Brought up server or GPU fleets at scale, hundreds of nodes or more, and taken them all the way to production.
- Work deep in Linux and out-of-band management: BMC, IPMI, and Redfish.
- Automated hardware workflows in Python or Go rather than clicking through them; turn repeated tasks into software.
- Worked physically in data halls, racking, cabling, and swapping components; effective acting as remote hands or directing them.
- Triage failures methodically across hardware, firmware, and software, isolating the fault to a component.
- Travel for turn-up windows when a new data hall comes online.
Nice-to-Haves
- Kubernetes-based bare-metal provisioning.
- Accelerator platform bringup (NVIDIA, AMD, or custom).
- Burn-in and stress harness design.
- DCIM and inventory tooling.