Skip to content
AnthropicAnthropic

Data Center Operations Lead - Partner Site Operations

Leads operations outcomes for partner-operated data center sites, directing vendors, defining operational standards, and ensuring deployment velocity, availability, repair performance, and incident response. Requires 8+ years in data center or infrastructure operations, vendor oversight experience, and hands-on server, network, and rack-level expertise.

About the job

Responsibilities

  • Own site availability, deployment milestones, and repair turnaround using independently verified data.
  • Set daily and weekly priorities for partner-operated data center sites and lead vendor operating cadences, including standups and business reviews.
  • Define and improve procedures for deployment, break-fix, change management, security, and EHS compliance.
  • Track vendor performance against SLAs and staffing commitments; drive corrective actions.
  • Participate in incident escalation on-call rotations and serve as Incident Commander for site-specific incidents.
  • Translate engineering requirements into vendor direction and communicate site constraints and risks to leadership.
  • Lead operations reviews and scorecards, deployment surges, root-cause analyses, post-mortems, readiness activities, and fleet-wide process improvements.

Requirements

  • 8+ years of experience in data center operations, hardware, IT infrastructure, or critical facilities, with accountability for production availability.
  • Experience managing vendors, MSPs, or contract workforces against SOWs, SLAs, operational reviews, and corrective actions.
  • Hands-on technical depth in server, network, and rack-level infrastructure.
  • Experience building or substantially improving operational processes.
  • Incident command or lead-responder experience and clear communication under ambiguity.
  • Ability to support non-standard hours, on-call rotations, deployment surges, and maintenance windows.
  • Bachelor's degree in a relevant field or equivalent practical experience.

Nice-to-haves

  • Experience with third-party colocation providers or partner-operated sites.
  • Experience standing up operations at a new site or data hall.
  • Experience with GPU or accelerator infrastructure and high-density liquid-cooled infrastructure.
  • Familiarity with multi-vendor sites where facilities and IT operations are handled by different partners.
  • Experience leading cross-functional projects without direct ownership of participating teams.
  • Background in incident management frameworks, contract/SLA design, or EHS programs.

Compensation

  • Annual salary: $320,000–$405,000 USD.

Skills

Data Center Operations, Hardware Infrastructure, It Infrastructure, Critical Facilities, Vendor Management, Service-Level Agreements, Server Infrastructure, Network Infrastructure, Rack-Level Infrastructure, Incident Management, Change Management, Root-Cause Analysis, Gpu Infrastructure, Liquid Cooling, Ehs Compliance

Vapi

Vapi

San Francisco, CA

Member of Technical Staff, Release Engineer
$235k+/yrHybrid7+ YOEDevOps / SRE

Own and improve the CI/CD, testing, and deployment infrastructure that enables fast, safe, observable releases at scale. The role requires strong distributed-systems expertise, hands-on Kubernetes and infrastructure-as-code experience, and a track record of measurable cross-team improvements.

The Voleon Group

The Voleon Group

Berkeley, CA
Senior Software Engineer, Developer Experience
$225k+/yrHybrid5+ YOEDevOps / SRE

Build and evolve the developer platform that enables reliable, efficient software delivery across the company. The role requires 5+ years of software engineering experience, strong programming and system-design fundamentals, and expertise in build systems, CI/CD, testing, and deployment automation.

Descript

Descript

San Francisco, CA

Software Engineer, Infrastructure
$220k+/yrRemote8+ YOEDevOps / SRE

Own and evolve a broad infrastructure platform spanning cloud, Kubernetes, deployment, reliability, security, and GPU-backed AI systems. The role requires 8+ years operating production distributed systems, strong incident and architecture experience, and practical cloud infrastructure expertise.

Zoox

Zoox

Foster City, CA

Senior Software Engineer - Pipeline Infrastructure & Integration
$219k+/yrHybrid7+ YOEDevOps / SRE

Senior engineer owning safety-critical software pipelines and infrastructure, from static and dynamic analysis through CI enforcement, dashboards, and reliability tooling. Requires an advanced technical degree, 7+ years working with large codebases, and expertise in Bazel, Python, backend infrastructure, and C++.

Skydio

Skydio

San Mateo, CA

Senior Software Engineer, Developer Productivity
$200k+/yrOn-site5+ YOEDevOps / SRE

Build and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.