# Staff Data Center Operations Engineer

**Company:** [Crusoe](https://hotfix.jobs/companies/crusoe)
**Location:** Denver, CO
**Role:** Hardware Engineering
**Salary:** $150k – $170k/yr
**Experience:** 7+ years
**Skills:** gpu infrastructure, data center operations, oem escalation, rma processes, server platform architecture, supermicro, hpe, bmc, pcie, thermal management, fabric, SOPs, runbooks, liquid cooling, amd instinct
**Posted:** 2026-07-15

> Serve as the senior technical operations expert for Crusoe's SiteOps organization, owning Tier 2/3 GPU hardware escalations, OEM/ODM partnerships, platform standards, SOPs, training, and cross-functional collaboration at HQ while traveling to sites as needed. Requires 7+ years of hands-on GPU data center operations experience.

## Job Description

## What You'll Do

### Cross-Site Platform Operations & Escalation
- Own Tier 2/3 hardware escalations across all Crusoe sites for issues that exceed local site capability, engaging directly with OEM and ODM engineering teams to drive resolution.
- Travel to sites as needed for complex platform issues, new hardware bring-ups, and deployment support.
- Identify recurring failure patterns across sites and translate them into platform feedback, sparing strategy inputs, or OEM improvement requests.
- Root-cause complex hardware issues — PCIe, BMC, thermal, fabric — and produce resolution documentation reusable across the SiteOps org.
- Hand off platform-level findings to the appropriate internal engineering teams with clear, well-documented escalation packages.

### OEM & ODM Technical Partnership
- Develop and maintain deep technical relationships with Crusoe's primary hardware partners — currently SuperMicro and HPE, with upcoming ODM’s as growing platforms — at the engineering and field escalation level.
- Serve as Crusoe's technical voice in OEM/ODM partner conversations, surfacing field observations, influencing hardware roadmaps, and driving platform improvements that benefit the full fleet.
- Build familiarity with new ODM platform architecture, tooling, and escalation processes as Crusoe expands its ODM footprint.
- Support vendor evaluations and new platform qualifications in partnership with SiteOps and engineering leadership.

### Platform Standards & Org Development
- Own OEM platform technical knowledge at the SiteOps org level — escalation playbooks, failure pattern analysis, and OEM relationship inputs across all sites.
- Own the development and maintenance of platform-specific SOPs, runbooks, and field troubleshooting procedures for the SiteOps org, ensuring site teams have current, actionable documentation across all active hardware platforms.
- Design and deliver technical training for SiteOps technicians covering hardware architecture, platform-specific troubleshooting, and field procedures — both for new hire onboarding and ongoing skill development as the fleet and team evolve.
- Contribute to the technician certification program and technical leveling standards across the org.
- Support new site bring-up efforts providing platform readiness and deployment execution expertise.

### HQ Presence & Cross-Functional Collaboration
- Serve as SiteOps' senior technical representative at Crusoe HQ, participating in platform, engineering, and procurement discussions that affect site operations.
- Partner with engineering and procurement teams on sparing strategy, RMA lifecycle management, and OEM/ODM support contract structures.
- Provide operational input into next-generation GPU platform evaluations (GB300, VR200, and beyond).
- Produce escalation reporting, platform health analysis, and operational insights for SiteOps leadership.

## What We're Looking For

### Required
- 7+ years in data center operations, field engineering, or OEM/ODM technical support with hands-on GPU infrastructure experience.
- Direct hands-on experience deploying and supporting GPU platforms at scale across one or more major OEMs or ODMs; familiarity with SuperMicro and HPE platforms required.
- Deep familiarity with server platform architecture and OEM escalation and RMA processes.
- Experience leading or contributing to large-scale GPU cluster bring-ups including rack staging and production handoff.
- Demonstrated ability to build technical relationships with OEM and ODM engineering teams and drive platform-level issue resolution.
- Experience developing SOPs, runbooks, or field troubleshooting procedures and delivering technical training to data center technician teams.
- Strong written communication — comfortable producing escalation documentation, platform runbooks, and leadership reporting.
- Willingness to travel domestically and internationally to Crusoe sites as needed (target: up to 30%).

### Preferred
- Direct experience with SuperMicro GPU platforms (B200, GB200, or newer); SuperMicro Certified Engineer credentials a plus.
- Familiarity with ASUS or Quanta server platforms and ODM engagement models.
- Experience with liquid-cooled GPU platforms and CDU integration.
- Familiarity with AMD Instinct GPU platforms (MI300X/MI350X/MI355X).
- Prior experience at an AI cloud provider, hyperscaler, or GPU-first infrastructure operator.
- Experience contributing to technician certification programs or IC leveling standards within a DC ops organization.

## Compensation
Compensation will be paid in the range of up to $150,000 - $170,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicants knowledge, education, and abilities, as well as internal equity and alignment with market data.

## Similar roles

- [3D Physical Design Engineer](https://hotfix.jobs/jobs/6d821cae-30ac-4b47-839a-bd1edaf8b1b3) - Cerebras Systems - Sunnyvale, CA - $150k – $270k/yr
- [AI Silicon Physical Design Engineer](https://hotfix.jobs/jobs/77169772-2289-4f2a-bd21-5daa2b1ae03e) - Cerebras Systems - Remote - $150k – $250k/yr
- [Staff Engineer, Electrical Failure Analysis (R4593)](https://hotfix.jobs/jobs/bc3ddc5d-1dc6-4ed3-96a1-0bdb698144c8) - Shield AI - Dallas, TX - $150k – $220k/yr
- [Staff Engineer, Power System Battery Pack Designer](https://hotfix.jobs/jobs/34f5670a-c7db-4787-8541-48a2d1bd0355) - Shield AI - Dallas, TX - $150k – $220k/yr
- [Senior Staff Engineer, Advanced Manufacturing (R4359)](https://hotfix.jobs/jobs/df856eb2-ef7b-41b3-88f0-a56e5ff3ecb9) - Shield AI - Dallas, TX - $150k – $230k/yr

**Apply:** https://hotfix.jobs/jobs/75c8cabd-b966-48fe-bb28-d1c97c497350
**Canonical:** https://hotfix.jobs/jobs/75c8cabd-b966-48fe-bb28-d1c97c497350