Skip to content
CrusoeCrusoe

Data Center Systems Engineer (MDC)

Leads integration of electrical, thermal, mechanical, and networking systems in modular data centers for high-density AI compute. Ensures compatibility with power sources, optimizes thermal management, and supports transition to liquid cooling with 6+ years systems engineering experience.

About the job

What You'll Be Working On

I. Physical-to-Digital Integration

  • System Convergence: Lead the integration of electrical, thermal, mechanical, and networking systems for the Spark MDC to support high-density compute requirements.
  • Power Architecture: Collaborate with Electrical Engineers to refine site one-line diagrams and internal power distribution, ensuring compatibility with grid-scale and microgrid power sources, including Spark-level back-up power systems.
  • Network Fabric: Partner with the Networking team to incorporate and standardize the deployment of leaf/spine switches and fiber management within the MDC, optimizing for both East-West (InfiniBand) and North-South connectivity. Ensure that connectivity of MDC to fiber etc is properly designed and accounted for.
  • Thermal Management: Partner and collaborate with CI (Crusoe Industries), & Product team to ensure internal thermal management of the MDC provides the proper environment to optimize compute/GPU performance and eliminate thermal throttling of any/all compute & networking functions.

II. Roadmap Evolution (Liquid Cooling)

  • Thermal Transition: Support the engineering transition from air-cooled to liquid-cooled (DLC/CDU) architectures, focusing on the systems-level impacts on power density and cooling efficiency.
  • Prototyping: Assist R&D in the testing and validation of next-gen "Spark" prototypes designed for next generation chip architectures.

III. Technical Standardization & Support

  • Manufacturability: Transition the Spark from a pilot-phase asset to a standardized, manufacturable product with documentation ready for contract manufacturing partners, including and incorporating best-in class DFM practices.
  • Tier 3 Support: Act as the technical escalation point for field teams during the commissioning of complex, multi-unit deployments.

What You'll Bring to the Team

  • Systems Engineering Mastery: 6+ years in systems engineering or data center design, with a track record of integrating complex power and mechanical systems.
  • Power & Thermal Fluency: Deep understanding of medium-voltage power distribution and industrial-scale cooling; specific experience with the transition to liquid-cooled (DLC) systems is highly preferred.
  • Networking Integration: Practical knowledge of high-speed data center networking topologies (InfiniBand/RoCE) and the physical layer requirements for large-scale GPU clusters.
  • Vertical Integration Mindset: Experience working in environments where hardware and software are co-developed, requiring frequent collaboration with firmware and cloud networking teams.
  • Compliance & Standards: Knowledge of UL/CE listing processes and NEC requirements for modular or containerized equipment.
  • Technical Communication: Ability to translate complex engineering tradeoffs into clear decision frameworks for non-technical leadership.
  • Mountaineer Spirit: A relentless focus on technical preparation, safety-first design, and mastery of integration tools.

Benefits

  • Competitive compensation
  • Restricted Stock Units
  • Paid time off & paid holidays
  • Comprehensive health, dental & vision insurance
  • Employer contributions to HSA account
  • Paid parental leave
  • Paid life insurance, short-term and long-term disability
  • Professional development & tuition reimbursement
  • Mental health & wellness support
  • Commuter benefits (parking & transit)
  • Cell phone stipend
  • 401(k) Retirement plan with company match up to 4% of salary
  • Volunteer time off

Compensation Range

Compensation will be paid in the range of up to $148,740 - $170,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicants knowledge, education, and abilities, as well as internal equity and alignment with market data.

Skills

InfiniBand, Roce, Leaf/Spine, Liquid Cooling, Dlc, Cdu, Medium-Voltage Power, Microgrid, Fiber Management, Nec, Ul/Ce

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

Gumloop

Gumloop

San Francisco, CA
Senior Infrastructure Engineer
$150k+/yrOn-siteDevOps / SRE

Own and scale infrastructure for agent orchestration, sandboxing, and hosted MCP services. The role requires hands-on Kubernetes, cloud, and infrastructure-as-code experience, along with strong software engineering fundamentals and high ownership.

Axle

Axle

Frederick, MD

IT Operations Technical Lead
$150k+/yrHybrid10+ YOEDevOps / SRE

Leads hybrid cloud and on-premises IT operations, incident management, automation, security hardening, and infrastructure reliability while mentoring systems engineers. Requires extensive Linux administration, ITIL operations, cloud migration, automation, and AI/ML infrastructure experience.

PointClickCare

PointClickCare

United States

Senior Site Reliability Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

Senior Site Reliability Engineer providing technical leadership for scalable operations, automation, monitoring, resiliency, and cloud infrastructure. Requires a bachelor's degree, software development or architecture experience, and hands-on DevOps or systems administration experience.

Okta

Okta

San Francisco, CA

Senior Site Reliability Engineer
$147k+/yrHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for operating and improving large-scale, FedRAMP-compliant cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, software engineering, and reliability engineering expertise.