Manager, Data Center Operations
Manages data center technicians and critical infrastructure supporting AI compute systems, including power, cooling, networking, hardware deployments, incidents, vendors, and capacity expansion. Requires 5+ years in data center operations and 3+ years managing technical teams.
About the job
Responsibilities
- Oversee power, cooling, networking, and hardware deployments to maintain 99.999% uptime for AI compute systems.
- Lead and develop Data Center Operations Technicians through training, performance evaluations, and team development.
- Manage hardware lifecycles, incident resolution, and inventory processes.
- Coordinate technicians, AI specialists, and external vendors during technology integrations and capacity expansions.
- Implement energy-efficient practices and sustainability initiatives.
- Track uptime, power efficiency, and issue-resolution metrics to improve site performance.
- Direct emergency response and resolve operational issues quickly.
- Develop preventative maintenance schedules with vendors and ticket workflows in Jira.
- Help standardize operational best practices across sites.
Requirements
- 5+ years of experience in data center operations or similar critical environments.
- 3+ years managing technical teams.
- Proven ability to lead teams in fast-paced, high-responsibility settings.
- Expertise in server hardware, cabling, and data center technologies, including setup and lifecycle management.
- Ability to work in a dynamic environment with occasional on-call duties.
- Willingness to travel to data center locations as needed.
- Ability to lift up to 50 lbs, stand for long periods, and occasionally use ladders.
Nice-to-Haves
- Experience supporting AI, machine learning, or high-performance computing environments.
- Proficiency with Jira and collaborative workflows.
- Strong analytical and technical communication skills.
- Familiarity with Python or Bash scripting for automation.
- Experience partnering with vendors, scaling operations, and advancing sustainability initiatives.
Skills
Server Hardware, Data Center Operations, Cabling, Networking, Power Management, Cooling Systems, Jira, Python, Bash, Incident Management, Inventory Management, Preventive Maintenance, High-Performance Computing, Machine Learning, Sustainability
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.