Senior Manager, Facilities Remote Operations Center
Leads Crusoe’s 24/7 Facility Remote Operations Center in Dallas, overseeing staff, monitoring systems, incident response, and operational processes across a distributed AI data center portfolio. Requires 5+ years in data center operations, MEP expertise, people leadership, and experience with critical-facility monitoring platforms.
About the job
Responsibilities
- Lead 24/7/365 Facility Remote Operations Center operations, including staffing, scheduling, shift coverage, and performance management.
- Establish and improve monitoring protocols, escalation procedures, and SOPs for facility alarms, events, and anomalies.
- Serve as the senior escalation point for critical facility events and coordinate responses with on-site teams, vendors, and leadership.
- Oversee monitoring of electrical, mechanical, fire/life-safety, and building automation systems across a distributed data center portfolio.
- Ensure DCIM, EPMS/BMS, ticketing, and alarm-management platforms are configured, integrated, and actionable.
- Own incident management, post-incident documentation, root cause analysis, after-action reviews, and corrective actions.
- Develop emergency response runbooks and escalation matrices for weather events, utility disturbances, and equipment failures.
- Coordinate with Data Center Operations, Critical Facilities Engineering, Construction/Commissioning, and Security teams.
- Support commissioning and turnover of new sites into ROC monitoring scope.
- Report ROC performance, uptime-impacting events, and trends to senior leadership.
- Recruit, train, and develop ROC staff, including technical competency and certification programs.
- Identify automation, tooling, and process-standardization opportunities to improve detection speed and reduce false-positive alarm fatigue.
Requirements
- 5+ years of experience in Data Center Operations with direct responsibility for critical facility uptime.
- Strong knowledge of data center MEP systems, including electrical distribution, generators, UPS, and air- and liquid-cooled mechanical systems in high-density AI/HPC environments.
- Experience in shift-based, 24/7 operations, including rotating shift scheduling.
- Experience with incident management, escalation procedures, and root cause analysis for critical facility events.
- Experience with DCIM, BMS/EPMS, or similar monitoring and alarm-management platforms.
- Proven people leadership experience, including hiring, coaching, and performance management.
- Strong communication skills for translating technical facility issues into actionable information.
- Ability to work on-site in Dallas, Texas, and support 24/7 operations, including off-hours escalations.
Nice-to-haves
- Hyperscale, colocation, or AI/GPU-cluster data center operations experience.
- CDCP, CDCS, CDCE, DCPRO, or equivalent industry certification.
- Experience establishing or scaling a remote or centralized operations function.
- Familiarity with liquid cooling infrastructure, including CDUs, manifolds, and rear-door heat exchangers.
- Experience supporting geographically distributed critical infrastructure portfolios.
- Bachelor's degree in Electrical Engineering, Mechanical Engineering, Facilities Management, or a related technical field, or equivalent practical experience.
Benefits
- Competitive compensation and equity packages
- Restricted Stock Units
- Paid time off, holidays, and leave programs
- Comprehensive health, dental, and vision insurance
- Employer HSA contributions
- Paid parental leave
- Paid life insurance and short- and long-term disability
- Professional development and tuition reimbursement
- Mental health and wellness support
- Commuter benefits
Skills
Data Center Operations, Mep Systems, Electrical Distribution, Generators, Ups, Mechanical Cooling, Liquid Cooling, Dcim, Bms, Epms, Incident Management, Root Cause Analysis, Alarm Management, Shift Scheduling, SOPs
Similar jobs
Engineering Management jobsLeads a remote-first team of SDETs responsible for product quality, test strategy, automation, and release reliability across SaaS and customer-managed environments. Requires 10+ years of industry experience, technical leadership, and strong expertise in testing, CI/CD, observability, and quality metrics.
Leads multiple service-infrastructure engineering teams and managers, shaping architecture, developer platforms, reliability, and cross-functional delivery. Requires substantial management experience, including managing managers, critical distributed systems, incident response, and geographically distributed teams.
Leads and develops the App Traffic engineering team building reliable, scalable service-mesh and networking infrastructure across multiple clouds. Requires 9+ years of software engineering experience, including engineering leadership and distributed-systems or infrastructure expertise.
Leads the mortgage engineering organization, owning platform architecture, delivery, business-line outcomes, and team development. Requires senior engineering management experience, extensive software engineering experience, large-team leadership, business ownership, and expertise in scalable systems and AI.
Leads a hands-on Shared Services Engineering team building and operating reusable services, SDKs, APIs, and customer-facing systems. Requires 7+ years of software engineering experience, engineering management experience, and strong technical judgment across distributed and full-stack systems.