What You'll Do
Cross-Site Platform Operations & Escalation
- Own Tier 2/3 hardware escalations across all Crusoe sites for issues that exceed local site capability, engaging directly with OEM and ODM engineering teams to drive resolution.
- Travel to sites as needed for complex platform issues, new hardware bring-ups, and deployment support.
- Identify recurring failure patterns across sites and translate them into platform feedback, sparing strategy inputs, or OEM improvement requests.
- Root-cause complex hardware issues — PCIe, BMC, thermal, fabric — and produce resolution documentation reusable across the SiteOps org.
- Hand off platform-level findings to the appropriate internal engineering teams with clear, well-documented escalation packages.
OEM & ODM Technical Partnership
- Develop and maintain deep technical relationships with Crusoe's primary hardware partners — currently SuperMicro and HPE, with upcoming ODM’s as growing platforms — at the engineering and field escalation level.
- Serve as Crusoe's technical voice in OEM/ODM partner conversations, surfacing field observations, influencing hardware roadmaps, and driving platform improvements that benefit the full fleet.
- Build familiarity with new ODM platform architecture, tooling, and escalation processes as Crusoe expands its ODM footprint.
- Support vendor evaluations and new platform qualifications in partnership with SiteOps and engineering leadership.
Platform Standards & Org Development
- Own OEM platform technical knowledge at the SiteOps org level — escalation playbooks, failure pattern analysis, and OEM relationship inputs across all sites.
- Own the development and maintenance of platform-specific SOPs, runbooks, and field troubleshooting procedures for the SiteOps org, ensuring site teams have current, actionable documentation across all active hardware platforms.
- Design and deliver technical training for SiteOps technicians covering hardware architecture, platform-specific troubleshooting, and field procedures — both for new hire onboarding and ongoing skill development as the fleet and team evolve.
- Contribute to the technician certification program and technical leveling standards across the org.
- Support new site bring-up efforts providing platform readiness and deployment execution expertise.
HQ Presence & Cross-Functional Collaboration
- Serve as SiteOps' senior technical representative at Crusoe HQ, participating in platform, engineering, and procurement discussions that affect site operations.
- Partner with engineering and procurement teams on sparing strategy, RMA lifecycle management, and OEM/ODM support contract structures.
- Provide operational input into next-generation GPU platform evaluations (GB300, VR200, and beyond).
- Produce escalation reporting, platform health analysis, and operational insights for SiteOps leadership.
What We're Looking For
Required
- 7+ years in data center operations, field engineering, or OEM/ODM technical support with hands-on GPU infrastructure experience.
- Direct hands-on experience deploying and supporting GPU platforms at scale across one or more major OEMs or ODMs; familiarity with SuperMicro and HPE platforms required.
- Deep familiarity with server platform architecture and OEM escalation and RMA processes.
- Experience leading or contributing to large-scale GPU cluster bring-ups including rack staging and production handoff.
- Demonstrated ability to build technical relationships with OEM and ODM engineering teams and drive platform-level issue resolution.
- Experience developing SOPs, runbooks, or field troubleshooting procedures and delivering technical training to data center technician teams.
- Strong written communication — comfortable producing escalation documentation, platform runbooks, and leadership reporting.
- Willingness to travel domestically and internationally to Crusoe sites as needed (target: up to 30%).
Preferred
- Direct experience with SuperMicro GPU platforms (B200, GB200, or newer); SuperMicro Certified Engineer credentials a plus.
- Familiarity with ASUS or Quanta server platforms and ODM engagement models.
- Experience with liquid-cooled GPU platforms and CDU integration.
- Familiarity with AMD Instinct GPU platforms (MI300X/MI350X/MI355X).
- Prior experience at an AI cloud provider, hyperscaler, or GPU-first infrastructure operator.
- Experience contributing to technician certification programs or IC leveling standards within a DC ops organization.
Compensation
Compensation will be paid in the range of up to $150,000 - $170,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicants knowledge, education, and abilities, as well as internal equity and alignment with market data.