Site Manager, Datacenter Operations
Lead 24/7 operations of critical data center infrastructure for AI compute sites, including electrical, mechanical, cooling, safety, hardware/network repair, incident response, and team leadership while sites are still under construction. Requires deep experience running contractual-availability facilities, technical infrastructure knowledge, and building high-performing teams.
About the job
Role Scope
- Lead 24/7/365 operation of all critical infrastructure at your site: electrical distribution (UPS, generators, switchgear), mechanical and liquid cooling, fire/life safety, and building controls, accountable for availability against customer SLAs.
- Own site safety in partnership with HSE, including hazardous energy control, electrical safety, and LOTO discipline across a mixed employee and contractor workforce.
- Build and lead the full site team, facility operators alongside hardware and network operators, owning shift coverage, on-call rotations, performance, and development.
- Run the hardware and network repair operation: own the ticket queues, drive triage, prioritization, break-fix, and RMA workflows, and hold time-to-repair against customer commitments as the fleet scales.
- Run incident and change management for the site: lead severity response as the escalation point, deliver root cause analyses on 5 and 10 day clocks, and drive corrective actions that fix classes of problems, not instances.
- Partner with construction, commissioning, colo providers, and utilities while the site is live and still being built, driving operational readiness, MOP and SOO rigor, and clean handover into operations.
What We're Looking For
- You've run 24/7 critical facility operations where availability was contractual, and you can walk through the incidents that tested it.
- You know electrical and mechanical infrastructure deeply enough to dig into the details with your technicians: you've stood in front of switchgear or a CDU and known which questions to ask.
- You've run a high-volume repair or break-fix operation, server fleets or network fabrics, where queue discipline and time-to-repair were the scoreboard.
- You've built and developed technical teams through hiring, coaching, and hard performance calls, and your former reports would say you made them better.
- You've operated procedure-based safety programs (LOTO, electrical safety, hazardous energy control) and held the line when schedule pressure pushed against them.
- You've partnered with construction and commissioning teams during live builds and caught operational and safety gaps before they reached handover.
- You write and present clearly enough to put an RCA or a status update in front of an executive or a customer without translation.
Bonus
- Hyperscale or Tier III/IV facility experience.
- Trade certification or journeyman license (electrical, HVAC, controls).
- New-facility commissioning and operational readiness programs.
- CMMS/EAM systems.
- GPU cluster or high-performance network operations background.
Salary & Benefits
- Competitive total compensation package (salary + equity).
- Retirement or pension plan, in line with local norms.
- Health, dental, and vision insurance.
- Generous PTO policy, in line with local norms.
- We are committed to pay equity and transparency.
Skills
Critical Infrastructure Operations, Electrical Distribution, Mechanical Cooling, Liquid Cooling, Fire/Life Safety Systems, Building Controls, Loto, Incident Management, Root Cause Analysis, Team Leadership, Hardware Repair Operations, Change Management, Data Center Commissioning
Similar jobs
Lead the build-out of a fleet-wide Enterprise Asset Management (EAM) and CMMS system from zero for a rapidly scaling data center operator. Own asset indexing, maintenance standards, program adoption, and data-driven reliability decisions across owned and colo sites.
Own external connectivity, backbone capacity planning, and carrier negotiations for multi-GW AI training networks across data center campuses. Requires scaled carrier/long-haul experience, dark fiber procurement, and capacity planning.
Leads the architecture, extensibility, integration, and AI strategy for Databricks’ SAP S/4HANA finance platform. Requires 10+ years in SAP technical roles, architecture or technical leadership experience, and delivery of a full S/4HANA program with BTP extensions.
Leads plant-level supplier quality response, containment, non-conformance disposition, and corrective actions to protect high-volume assembly operations. Requires a bachelor's degree and 6+ years of plant, manufacturing, or supplier quality experience, plus expertise in structured problem solving and core quality tools.
Lead AV Engineer owning multi-site audiovisual standards, infrastructure, event production, and the Zoom-to-Google Meet migration. The role requires 8+ years of AV engineering experience, deep control/audio/video expertise, and strong live-event and vendor leadership.