Skip to content
xAIxAI

Manager, Operations

Leads facilities operations, power generation, and fiber teams for xAI's hyperscale AI compute facilities, ensuring 24/7 uptime, efficiency, and reliability of critical infrastructure. Requires 5+ years in data center operations with management experience and deep expertise in power systems and networking.

About the job

Responsibilities

  • Lead and scale the facilities operations and power generation teams responsible for the reliable operation, maintenance, monitoring, and optimization of critical infrastructure including on-site power generation assets, electrical systems, mechanical/HVAC, liquid cooling, power distribution, UPS, generators, and building management systems.
  • Direct the fiber teams overseeing the design, deployment, maintenance, and expansion of high-speed fiber optic networks, dark fiber, and connectivity infrastructure supporting AI compute clusters and data center interconnects.
  • Own key performance metrics such as uptime (targeting 99.999%+), mean time to detect/repair (MTTD/MTTR), power usage effectiveness (PUE), water usage effectiveness (WUE), power generation efficiency, and overall infrastructure availability.
  • Develop and enforce standard operating procedures (SOPs), preventive maintenance programs, incident response protocols, and continuous improvement processes for both facilities and power generation assets to minimize downtime and maximize efficiency.
  • Build, mentor, and grow multidisciplinary teams of operations technicians, power generation engineers and controls specialists while fostering a culture of ownership, safety, and excellence.
  • Partner closely with engineering, construction, procurement, and AI hardware teams to support new facility builds, expansions, commissioning, power integration, and smooth handovers from project to operations.
  • Manage operational budgets, vendor relationships (maintenance contractors, fiber providers, power generation OEMs, fuel suppliers), spare parts inventory, and risk mitigation strategies in a high-velocity environment.
  • Drive innovation in operational practices, automation, predictive maintenance, power generation optimization, and sustainability initiatives to support the extreme power and cooling demands of next-generation AI systems.
  • Provide regular performance reporting, root cause analyses, lessons learned, and strategic recommendations to senior leadership.

Basic Qualifications

  • 5+ years of progressive experience in data center facilities operations, power generation operations, hyperscale infrastructure management, or mission-critical industrial operations, with at least 2+ years in a management or supervisor role.
  • Proven track record leading large-scale operations teams supporting high-density compute environments with significant on-site or dedicated power generation (AI, HPC, or hyperscaler data centers strongly preferred).
  • Strong experience managing fiber optic networks, dark fiber deployments, or high-bandwidth connectivity infrastructure in large-scale technical environments.
  • Deep knowledge of power generation systems (gas turbines, reciprocating engines, cogeneration), MEP (mechanical, electrical, plumbing) systems, BMS/SCADA, liquid cooling, power redundancy topologies, and 24/7 operations best practices.
  • Demonstrated success delivering high reliability, rapid incident resolution, and operational excellence under aggressive scaling timelines.
  • Hands-on leadership style with the ability to roll up sleeves while effectively managing teams, budgets, and cross-functional stakeholders.
  • Proficiency with operations tools, CMMS (computerized maintenance management systems), monitoring platforms, and data-driven decision making.

Preferred Skills and Experience

  • Direct background in AI or hyperscale data center operations, including liquid cooling systems, high-power GPU/accelerator environments, and on-site power generation.
  • Experience building or scaling fiber infrastructure for low-latency, high-bandwidth interconnects between compute clusters or sites.
  • Familiarity with Uptime Institute Tier standards, ASHRAE guidelines, power generation standards (e.g., IEEE, NFPA), OSHA/EPA compliance, and sustainability practices in critical facilities.
  • Track record of implementing automation, predictive analytics, or process improvements that significantly enhanced operational performance and power reliability.

Skills

Data Center Operations, Power Generation, Fiber Optic Networks, Mep Systems, Bms/Scada, Liquid Cooling, Cmms, Gas Turbines, Predictive Maintenance, Automation

Similar jobs

HumanSignal

HumanSignal

Columbus, OH

Paid Voice Acting & Improv Work
$104k+/yrOn-siteOther

Performs unscripted household scenarios and natural conversations with a voice assistant during in-person AI training recordings. Requires improv or performance experience, comfort with camera and audio recording, and reliable attendance; compensation is $50 per hour.

Mercury

Mercury

San Francisco, CA
AML Investigator III
$110k+/yrRemoteOther

Investigate AML, fraud, and other financial-crime alerts; respond to partner-bank information requests, assess customer activity, and escalate suspicious behavior. The role requires financial-services investigation experience, transaction-monitoring expertise, and strong communication and organizational skills.

Celonis

Celonis

New York, NY

US Only Job Description Template
No salary listedHybridOther

This posting is an incomplete role template for a full-time hybrid position in New York. Responsibilities and qualifications have not been provided; the posting highlights Celonis’s process intelligence platform and employee benefits.

Ascertain

Ascertain

United States

Prior Authorization Assistant, Cardiology
$45k+/yrRemote4+ YOEOther

Processes and oversees cardiology prior authorizations, coordinates documentation with healthcare providers and payors, and helps improve AI-powered healthcare workflows. Requires 4+ years of cardiology prior authorization and billing experience plus athenahealth or NextGen proficiency.

OpenAI

OpenAI

San Francisco, CA

Child Safety Enforcement Specialist
$158k+/yrHybridOther

Leads child safety investigations, enforcement decisions, mandatory-reporting workflows, and process improvements involving sensitive content and abuse signals. The role requires Trust & Safety investigation experience, sound judgment, strong documentation, and cross-functional collaboration.