Data Center Operations Lead - Partner Site Operations
Leads operations outcomes for partner-operated data center sites, directing vendors, defining operational standards, and ensuring deployment velocity, availability, repair performance, and incident response. Requires 8+ years in data center or infrastructure operations, vendor oversight experience, and hands-on server, network, and rack-level expertise.
About the job
Responsibilities
- Own site availability, deployment milestones, and repair turnaround using independently verified data.
- Set daily and weekly priorities for partner-operated data center sites and lead vendor operating cadences, including standups and business reviews.
- Define and improve procedures for deployment, break-fix, change management, security, and EHS compliance.
- Track vendor performance against SLAs and staffing commitments; drive corrective actions.
- Participate in incident escalation on-call rotations and serve as Incident Commander for site-specific incidents.
- Translate engineering requirements into vendor direction and communicate site constraints and risks to leadership.
- Lead operations reviews and scorecards, deployment surges, root-cause analyses, post-mortems, readiness activities, and fleet-wide process improvements.
Requirements
- 8+ years of experience in data center operations, hardware, IT infrastructure, or critical facilities, with accountability for production availability.
- Experience managing vendors, MSPs, or contract workforces against SOWs, SLAs, operational reviews, and corrective actions.
- Hands-on technical depth in server, network, and rack-level infrastructure.
- Experience building or substantially improving operational processes.
- Incident command or lead-responder experience and clear communication under ambiguity.
- Ability to support non-standard hours, on-call rotations, deployment surges, and maintenance windows.
- Bachelor's degree in a relevant field or equivalent practical experience.
Nice-to-haves
- Experience with third-party colocation providers or partner-operated sites.
- Experience standing up operations at a new site or data hall.
- Experience with GPU or accelerator infrastructure and high-density liquid-cooled infrastructure.
- Familiarity with multi-vendor sites where facilities and IT operations are handled by different partners.
- Experience leading cross-functional projects without direct ownership of participating teams.
- Background in incident management frameworks, contract/SLA design, or EHS programs.
Compensation
- Annual salary: $320,000–$405,000 USD.
Skills
Data Center Operations, Hardware Infrastructure, It Infrastructure, Critical Facilities, Vendor Management, Service-Level Agreements, Server Infrastructure, Network Infrastructure, Rack-Level Infrastructure, Incident Management, Change Management, Root-Cause Analysis, Gpu Infrastructure, Liquid Cooling, Ehs Compliance
Similar jobs
DevOps / SRE jobsOwn and improve the CI/CD, testing, and deployment infrastructure that enables fast, safe, observable releases at scale. The role requires strong distributed-systems expertise, hands-on Kubernetes and infrastructure-as-code experience, and a track record of measurable cross-team improvements.
Build and evolve the developer platform that enables reliable, efficient software delivery across the company. The role requires 5+ years of software engineering experience, strong programming and system-design fundamentals, and expertise in build systems, CI/CD, testing, and deployment automation.
Own and evolve a broad infrastructure platform spanning cloud, Kubernetes, deployment, reliability, security, and GPU-backed AI systems. The role requires 8+ years operating production distributed systems, strong incident and architecture experience, and practical cloud infrastructure expertise.
Senior engineer owning safety-critical software pipelines and infrastructure, from static and dynamic analysis through CI enforcement, dashboards, and reliability tooling. Requires an advanced technical degree, 7+ years working with large codebases, and expertise in Bazel, Python, backend infrastructure, and C++.
Build and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.