CPU/Storage/PoP-WAN Program Manager
Leads execution of CPU, storage, PoP, and WAN infrastructure programs to activate compute clusters and expand global networks. Requires 8+ years in technical program management with deep knowledge of hardware, networking, and data center deployments.
About the job
Key Responsibilities
- Lead end-to-end execution of CPU / GPU cluster activation programs across OpenAI’s global infrastructure footprint
- Drive readiness to convert contracted compute capacity into schedulable production clusters
- Own deployment programs for new PoPs, backbone nodes, WAN expansion, and interconnection initiatives
- Build integrated schedules spanning procurement, logistics, installation, storage readiness, network turn-up, testing, and production handoff
- Coordinate BOM readiness, server delivery, racks, optics, cabling, storage hardware, and vendor milestones
- Partner with engineering teams to align compute, storage, and networking dependencies before cluster activation
- Manage deployment of storage systems supporting training and inference workloads, including readiness, validation, performance checks, and scaling plans
- Coordinate backbone capacity expansion, cross-connects, inter-region pathing, and cloud interconnect readiness with Azure and third-party providers
- Lead physical deployment execution including rack-and-stack, hardware bring-up, L1 validation, and site acceptance criteria
- Build repeatable deployment playbooks, dashboards, governance cadences, and operating mechanisms for scale
- Identify risks early across supply chain, site readiness, technical constraints, and vendor execution, then drive mitigation plans
- Communicate milestones, escalations, and capacity forecasts to senior leadership
Qualifications
- 8+ years of experience in technical program management, infrastructure deployment, network deployment, or data center operations
- Strong experience delivering programs involving compute, storage, networking, or large-scale infrastructure systems
- Working knowledge of servers, clusters, storage arrays, routers, switches, optics, and structured cabling
- Experience owning cross-functional programs across engineering, operations, supply chain, and external vendors
- Strong understanding of deployment lifecycles from planning and procurement through production handoff
- Ability to reason across physical infrastructure execution and logical systems architecture dependencies
- Proven ability to build integrated schedules and drive accountability across multiple stakeholders
- Strong executive communication skills with experience managing critical escalations and leadership updates
- Comfortable operating in fast-moving environments with aggressive timelines and evolving priorities
- Highly analytical with strong problem-solving and execution instincts
Preferred Skills
- Experience at a hyperscaler, cloud provider, AI infrastructure company, or global network operator
- Experience deploying GPU clusters, HPC systems, or large training environments
- Familiarity with distributed storage systems and high-performance data infrastructure
- Experience with PoP deployments, WAN backbone expansion, or global network buildouts
- Experience working across first-party, colo, and cloud environments
- Experience building repeatable infrastructure deployment systems in high-growth environments
Skills
Gpu Clusters, Storage Systems, Pop Deployments, Wan, Data Center Operations, Servers, Routers, Switches, Optics, Structured Cabling, Hpc Systems, Distributed Storage, Cloud Interconnects, Azure
Similar jobs
DevOps / SRE jobsLeads operations outcomes for partner-operated data center sites, directing vendors, defining operational standards, and ensuring deployment velocity, availability, repair performance, and incident response. Requires 8+ years in data center or infrastructure operations, vendor oversight experience, and hands-on server, network, and rack-level expertise.
Own and improve the CI/CD, testing, and deployment infrastructure that enables fast, safe, observable releases at scale. The role requires strong distributed-systems expertise, hands-on Kubernetes and infrastructure-as-code experience, and a track record of measurable cross-team improvements.
Build and evolve the developer platform that enables reliable, efficient software delivery across the company. The role requires 5+ years of software engineering experience, strong programming and system-design fundamentals, and expertise in build systems, CI/CD, testing, and deployment automation.
Own and evolve a broad infrastructure platform spanning cloud, Kubernetes, deployment, reliability, security, and GPU-backed AI systems. The role requires 8+ years operating production distributed systems, strong incident and architecture experience, and practical cloud infrastructure expertise.
Senior engineer owning safety-critical software pipelines and infrastructure, from static and dynamic analysis through CI enforcement, dashboards, and reliability tooling. Requires an advanced technical degree, 7+ years working with large codebases, and expertise in Bazel, Python, backend infrastructure, and C++.