Senior Manager, Data Center Facility Operations
Leads 24/7 critical facility operations for a 20 MW AI data center expanding to 40 MW, overseeing electrical, mechanical, safety, maintenance, staffing, and operational performance. Requires at least five years of data center operations management experience and expertise in infrastructure reliability and team scaling.
About the job
Responsibilities
Critical Facility Operations
- Own 24/7 operational readiness for electrical distribution systems, including switchgear, UPS systems, generators, and PDUs.
- Oversee mechanical systems such as CRAH/CRAC units, chillers, cooling towers, and liquid/direct-to-chip cooling where applicable.
- Manage fire/life-safety systems and BMS/EPMS monitoring platforms.
- Maintain and improve Methods of Procedure (MOPs), Emergency Operating Procedures (EOPs), and Standard Operating Procedures (SOPs).
- Lead incident root-cause analysis, corrective actions, postmortems, and remediation tracking.
- Manage preventive and predictive maintenance programs to maximize reliability and minimize unplanned downtime.
Team and Staffing
- Build and manage a lean, cross-trained operations team of facility engineers and technicians.
- Define staffing ratios, shift structures, on-call coverage, and scaling plans for expansion from 20 MW to 40 MW.
- Recruit, develop, retain, and cross-train technical personnel.
- Manage subcontractors, vendors, Critical Facility Managers, and staffing partners.
Capacity Expansion
- Represent operations in expansion planning with Design, Construction, and Commissioning teams.
- Integrate new capacity without disrupting live production loads.
- Support integrated systems testing and commissioning of new electrical and mechanical infrastructure.
- Update staffing plans, spares strategies, and maintenance programs for the expanded footprint.
Safety, Compliance, and Reporting
- Champion a zero-incident safety culture and ensure compliance with OSHA, NFPA 70E, arc-flash requirements, LOTO procedures, local Minnesota and Shakopee codes, environmental permits, and utility interconnection agreements.
- Own KPIs including uptime, availability, PUE, MTTR, MTBF, preventive-maintenance completion, and safety metrics.
- Provide operational reporting to regional and corporate leadership.
- Manage site OpEx budgets covering labor, spares, contracts, and utilities.
Requirements
- At least 5 years managing operations in a data center environment with direct responsibility for critical electrical and mechanical infrastructure.
- Experience building, right-sizing, or optimizing lean staffing models, including shift design, cross-training, and vendor balance.
- Strong knowledge of electrical distribution and mechanical cooling systems in enterprise, hyperscale, or colocation data centers.
- Experience managing direct employees and/or contracted staff in a 24/7 operational environment.
- Working knowledge of BMS, EPMS, or DCIM monitoring platforms.
- Familiarity with NFPA 70E, OSHA, and industry-standard safety programs.
- Strong incident-management and root-cause-analysis skills.
- Bachelor's degree in Electrical Engineering, Mechanical Engineering, Facilities Management, or equivalent experience.
- Ability to work onsite, participate in an on-call rotation, and respond onsite during critical incidents.
Nice-to-Haves
- Experience with high-density AI/GPU compute environments, including liquid or direct-to-chip cooling.
- Experience operating during live capacity expansions or brownfield build-outs.
- Uptime Institute ATD/ADOS, CDCP/CDCS, or similar data center credentials.
- Experience with Upper Midwest utility and regulatory environments; Xcel Energy service-territory knowledge is a plus.
Compensation and Benefits
- Compensation range up to $175,000–$190,000 plus bonus.
- Restricted Stock Units included in all offers.
- Paid time off, paid holidays, and leave programs.
- Health, dental, and vision insurance.
- Employer HSA contributions.
- Paid parental leave.
- Life insurance and short- and long-term disability coverage.
- Professional development and tuition reimbursement.
- Mental health and wellness support.
- Commuter benefits and cell phone stipend.
- 401(k) plan with company match up to 4% of salary.
- Volunteer time off, global travel insurance, meals allowance, and location-specific perks.
Skills
Electrical Distribution, Mechanical Cooling, Bms, Epms, Dcim, Nfpa 70E, Osha, Root Cause Analysis, Preventive Maintenance, Liquid Cooling, Commissioning, Incident Management, Loto
Similar jobs
Engineering Management jobsLeads the Autonomous Behaviors and Agents engineering team, setting technical direction and delivering real-time multi-agent planning and optimization systems for mission-critical autonomy. Requires substantial modern C++ experience, algorithmic depth, and demonstrated engineering leadership.
Leads Snowflake Professional Services delivery, managing Solutions Architects and Consultants while overseeing complex data migrations and AI or generative AI engagements. Requires 10+ years in customer-facing technical roles, 5+ years of people management, and substantial professional-services sales experience.
Leads the Config Deployments engineering team responsible for high-scale configuration distribution, feature flags, and service coordination infrastructure. The role requires 7+ years of engineering experience, infrastructure expertise, and engineering management experience.
Leads the architecture and development of BuildOps’ integration and data platform, driving technical strategy, reliability, APIs, databases, and engineering standards across multiple teams. Requires 10+ years of software engineering experience and deep expertise in TypeScript, Node.js, PostgreSQL, cloud infrastructure, and distributed systems.
Leads a hands-on Shared Services Engineering team building and operating reusable services, SDKs, APIs, and customer-facing systems. Requires 7+ years of software engineering experience, engineering management experience, and strong technical judgment across distributed and full-stack systems.