Skip to content
CrusoeCrusoeDenver, CO

Staff Data Center Operations Engineer

Serve as the senior technical operations expert for Crusoe's SiteOps organization, owning Tier 2/3 GPU hardware escalations, OEM/ODM partnerships, platform standards, SOPs, training, and cross-functional collaboration at HQ while traveling to sites as needed. Requires 7+ years of hands-on GPU data center operations experience.

150k – 170k/yr
On-site7+ YOEHardware Engineering

About the role

What You'll Do

Cross-Site Platform Operations & Escalation

  • Own Tier 2/3 hardware escalations across all Crusoe sites for issues that exceed local site capability, engaging directly with OEM and ODM engineering teams to drive resolution.
  • Travel to sites as needed for complex platform issues, new hardware bring-ups, and deployment support.
  • Identify recurring failure patterns across sites and translate them into platform feedback, sparing strategy inputs, or OEM improvement requests.
  • Root-cause complex hardware issues — PCIe, BMC, thermal, fabric — and produce resolution documentation reusable across the SiteOps org.
  • Hand off platform-level findings to the appropriate internal engineering teams with clear, well-documented escalation packages.

OEM & ODM Technical Partnership

  • Develop and maintain deep technical relationships with Crusoe's primary hardware partners — currently SuperMicro and HPE, with upcoming ODM’s as growing platforms — at the engineering and field escalation level.
  • Serve as Crusoe's technical voice in OEM/ODM partner conversations, surfacing field observations, influencing hardware roadmaps, and driving platform improvements that benefit the full fleet.
  • Build familiarity with new ODM platform architecture, tooling, and escalation processes as Crusoe expands its ODM footprint.
  • Support vendor evaluations and new platform qualifications in partnership with SiteOps and engineering leadership.

Platform Standards & Org Development

  • Own OEM platform technical knowledge at the SiteOps org level — escalation playbooks, failure pattern analysis, and OEM relationship inputs across all sites.
  • Own the development and maintenance of platform-specific SOPs, runbooks, and field troubleshooting procedures for the SiteOps org, ensuring site teams have current, actionable documentation across all active hardware platforms.
  • Design and deliver technical training for SiteOps technicians covering hardware architecture, platform-specific troubleshooting, and field procedures — both for new hire onboarding and ongoing skill development as the fleet and team evolve.
  • Contribute to the technician certification program and technical leveling standards across the org.
  • Support new site bring-up efforts providing platform readiness and deployment execution expertise.

HQ Presence & Cross-Functional Collaboration

  • Serve as SiteOps' senior technical representative at Crusoe HQ, participating in platform, engineering, and procurement discussions that affect site operations.
  • Partner with engineering and procurement teams on sparing strategy, RMA lifecycle management, and OEM/ODM support contract structures.
  • Provide operational input into next-generation GPU platform evaluations (GB300, VR200, and beyond).
  • Produce escalation reporting, platform health analysis, and operational insights for SiteOps leadership.

What We're Looking For

Required

  • 7+ years in data center operations, field engineering, or OEM/ODM technical support with hands-on GPU infrastructure experience.
  • Direct hands-on experience deploying and supporting GPU platforms at scale across one or more major OEMs or ODMs; familiarity with SuperMicro and HPE platforms required.
  • Deep familiarity with server platform architecture and OEM escalation and RMA processes.
  • Experience leading or contributing to large-scale GPU cluster bring-ups including rack staging and production handoff.
  • Demonstrated ability to build technical relationships with OEM and ODM engineering teams and drive platform-level issue resolution.
  • Experience developing SOPs, runbooks, or field troubleshooting procedures and delivering technical training to data center technician teams.
  • Strong written communication — comfortable producing escalation documentation, platform runbooks, and leadership reporting.
  • Willingness to travel domestically and internationally to Crusoe sites as needed (target: up to 30%).

Preferred

  • Direct experience with SuperMicro GPU platforms (B200, GB200, or newer); SuperMicro Certified Engineer credentials a plus.
  • Familiarity with ASUS or Quanta server platforms and ODM engagement models.
  • Experience with liquid-cooled GPU platforms and CDU integration.
  • Familiarity with AMD Instinct GPU platforms (MI300X/MI350X/MI355X).
  • Prior experience at an AI cloud provider, hyperscaler, or GPU-first infrastructure operator.
  • Experience contributing to technician certification programs or IC leveling standards within a DC ops organization.

Compensation

Compensation will be paid in the range of up to $150,000 - $170,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicants knowledge, education, and abilities, as well as internal equity and alignment with market data.

Skills

gpu infrastructuredata center operationsoem escalationrma processesserver platform architecturesupermicrohpebmcpciethermal managementfabricSOPsrunbooksliquid coolingamd instinct
Cerebras Systems

3D Physical Design Engineer

Cerebras SystemsSunnyvale, CA

Designs and analyzes 3D integrated ASIC/SoC products, optimizing power, performance, area, and handling verification, IR/EM, packaging, and cooling. Requires 10+ years experience in physical design flows, scripting, and 3D stacking technologies.

150k – 270k/yr
On-site10+ YOEHardware Engineering
Cerebras Systems

AI Silicon Physical Design Engineer

Cerebras SystemsSunnyvale, CA +1

Performs physical design for AI silicon chips, including synthesis, place & route, timing closure, and verification for high-speed wafer-scale architectures. Requires 10+ years experience with EDA tools and scripting.

150k – 250k/yr
Remote10+ YOEHardware Engineering
Shield AI

Staff Engineer, Electrical Failure Analysis (R4593)

Shield AIDallas, TX

Leads failure analysis investigations for avionics hardware including PCBAs, harnesses, and interconnects. Performs root cause analysis using structured methods and drives corrective actions. Requires 8+ years in aerospace/defense hardware reliability and hands-on FA techniques.

150k – 220k/yr
On-site8+ YOEHardware Engineering
Shield AI

Staff Engineer, Power System Battery Pack Designer

Shield AIDallas, TX +2

Lead battery system architecture, design, qualification, and integration for high-voltage (270V+) power systems supporting 100kW applications in harsh environments. Requires 5+ years delivering battery systems to market and a BS in Electrical Engineering.

150k – 220k/yr
On-site5+ YOEHardware Engineering
Shield AI

Senior Staff Engineer, Advanced Manufacturing (R4359)

Shield AIDallas, TX

Develops manufacturing processes, ergonomic workstations, and work instructions for aerospace assemblies to drive production stability and rate acceleration. Requires 7+ years in manufacturing engineering, bachelor's in STEM, and experience with prototype production in fast-paced environments.

150k – 230k/yr
On-site7+ YOEHardware Engineering