Staff Data Center Operations Engineer
Serve as the senior technical operations expert for Crusoe's SiteOps organization, owning Tier 2/3 GPU hardware escalations, OEM/ODM partnerships, platform standards, SOPs, training, and cross-functional collaboration at HQ while traveling to sites as needed. Requires 7+ years of hands-on GPU data center operations experience.
About the job
What You'll Do
Cross-Site Platform Operations & Escalation
- Own Tier 2/3 hardware escalations across all Crusoe sites for issues that exceed local site capability, engaging directly with OEM and ODM engineering teams to drive resolution.
- Travel to sites as needed for complex platform issues, new hardware bring-ups, and deployment support.
- Identify recurring failure patterns across sites and translate them into platform feedback, sparing strategy inputs, or OEM improvement requests.
- Root-cause complex hardware issues — PCIe, BMC, thermal, fabric — and produce resolution documentation reusable across the SiteOps org.
- Hand off platform-level findings to the appropriate internal engineering teams with clear, well-documented escalation packages.
OEM & ODM Technical Partnership
- Develop and maintain deep technical relationships with Crusoe's primary hardware partners — currently SuperMicro and HPE, with upcoming ODM’s as growing platforms — at the engineering and field escalation level.
- Serve as Crusoe's technical voice in OEM/ODM partner conversations, surfacing field observations, influencing hardware roadmaps, and driving platform improvements that benefit the full fleet.
- Build familiarity with new ODM platform architecture, tooling, and escalation processes as Crusoe expands its ODM footprint.
- Support vendor evaluations and new platform qualifications in partnership with SiteOps and engineering leadership.
Platform Standards & Org Development
- Own OEM platform technical knowledge at the SiteOps org level — escalation playbooks, failure pattern analysis, and OEM relationship inputs across all sites.
- Own the development and maintenance of platform-specific SOPs, runbooks, and field troubleshooting procedures for the SiteOps org, ensuring site teams have current, actionable documentation across all active hardware platforms.
- Design and deliver technical training for SiteOps technicians covering hardware architecture, platform-specific troubleshooting, and field procedures — both for new hire onboarding and ongoing skill development as the fleet and team evolve.
- Contribute to the technician certification program and technical leveling standards across the org.
- Support new site bring-up efforts providing platform readiness and deployment execution expertise.
HQ Presence & Cross-Functional Collaboration
- Serve as SiteOps' senior technical representative at Crusoe HQ, participating in platform, engineering, and procurement discussions that affect site operations.
- Partner with engineering and procurement teams on sparing strategy, RMA lifecycle management, and OEM/ODM support contract structures.
- Provide operational input into next-generation GPU platform evaluations (GB300, VR200, and beyond).
- Produce escalation reporting, platform health analysis, and operational insights for SiteOps leadership.
What We're Looking For
Required
- 7+ years in data center operations, field engineering, or OEM/ODM technical support with hands-on GPU infrastructure experience.
- Direct hands-on experience deploying and supporting GPU platforms at scale across one or more major OEMs or ODMs; familiarity with SuperMicro and HPE platforms required.
- Deep familiarity with server platform architecture and OEM escalation and RMA processes.
- Experience leading or contributing to large-scale GPU cluster bring-ups including rack staging and production handoff.
- Demonstrated ability to build technical relationships with OEM and ODM engineering teams and drive platform-level issue resolution.
- Experience developing SOPs, runbooks, or field troubleshooting procedures and delivering technical training to data center technician teams.
- Strong written communication — comfortable producing escalation documentation, platform runbooks, and leadership reporting.
- Willingness to travel domestically and internationally to Crusoe sites as needed (target: up to 30%).
Preferred
- Direct experience with SuperMicro GPU platforms (B200, GB200, or newer); SuperMicro Certified Engineer credentials a plus.
- Familiarity with ASUS or Quanta server platforms and ODM engagement models.
- Experience with liquid-cooled GPU platforms and CDU integration.
- Familiarity with AMD Instinct GPU platforms (MI300X/MI350X/MI355X).
- Prior experience at an AI cloud provider, hyperscaler, or GPU-first infrastructure operator.
- Experience contributing to technician certification programs or IC leveling standards within a DC ops organization.
Compensation
Compensation will be paid in the range of up to $150,000 - $170,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicants knowledge, education, and abilities, as well as internal equity and alignment with market data.
Skills
Gpu Infrastructure, Data Center Operations, Oem Escalation, Rma Processes, Server Platform Architecture, Supermicro, Hpe, Bmc, Pcie, Thermal Management, Fabric, SOPs, Runbooks, Liquid Cooling, Amd Instinct
Similar jobs
Hardware Engineering jobsLeads flight-critical high- and low-voltage power electronics design for an autonomous aircraft, from concept through production and validation. Requires a bachelor’s degree in electrical engineering, 6+ years of power electronics experience, product-shipping experience, and technical leadership.
Leads networking, systems administration, instrumentation, and controls infrastructure for mineral-refining facilities, from equipment design through commissioning and ramp. Requires at least seven years of control-system networking experience, expert Ignition knowledge, and experience delivering large capital projects.
Leads instrumentation and controls engineering for mineral-refining equipment and plant infrastructure from design through commissioning and operations handover. Requires 7+ years of instrumentation experience, control and HMI programming, capital-project execution, and leadership of engineering teams.
Leads transistor architecture, electrical characterization, parameter extraction, and diagnostic analysis for in-house semiconductor device development. Requires a Ph.D. and deep expertise in CMOS/MOSFET physics, semiconductor measurement, probe stations, and data analysis.
Leads manufacturing engineering for advanced propulsion and fluid-system hardware from development through production, including process development, integration, validation, and production support. Requires 7+ years of relevant experience, a bachelor's degree, and expertise in aerospace propulsion manufacturing and DFM/DFA.