Technical Program Manager, Compute Infrastructure
Leads end-to-end delivery of large-scale GPU clusters and compute infrastructure, managing hardware, networking, power, cooling across partners. Requires 5+ years program management in hyperscaler infrastructure and technical expertise partnering with engineering teams.
About the job
In this role, you will:
- Lead end-to-end delivery of both New Compute SKUs and large-scale GPU clusters across an external partner ecosystem while supporting capacity planning for training and inference.
- Ability to contextually drive multi-threaded bring-up programs spanning hardware, networking, power, and cooling—owning plans, dependencies, and critical paths.
- Interface with chip providers to derisk long-term onboarding to new hardware platforms by working across kernels, comms, hardware, and scheduling engineering teams.
- Build and operationalize program mechanisms (roadmaps, milestones, risk registers, runbooks) that make delivery predictable at massive scale.
- Partner with engineering to improve cluster turn-up reliability, repeatability, and automation, reducing time-to-serve for new capacity.
- Coordinate cross-functional readiness (security, finance, operations, product/research stakeholders) to ship production-ready compute.
- Manage integration and handoffs across teams and partners—ensuring consistent execution, clear communication, and fast issue resolution.
- Identify bottlenecks and systemic gaps, then drive durable fixes across tooling, process, and partner interfaces.
- Provide crisp executive visibility on progress, tradeoffs, and risks across a large portfolio of concurrent programs.
You might thrive in this role if you:
- Possess a degree in a hard science, or have a demonstrated track record of engineering expertise.
- Have 5+ years of experience in program management for major projects including capital projects or hyperscaler infrastructure deployment.
- Demonstrate the ability to serve as the go-to person solely responsible for driving and delivering complex projects.
- Are comfortable managing cross-functional and cross-company teams; experience driving information and decision hygiene.
- Have an extensive track record of successfully delivering high-profile, technical projects against tight deadlines.
- Are technically adept and have effectively partnered with engineering or fundamental research teams of the highest caliber.
- Have experience interfacing with and leading external vendors including engineering firms, equipment suppliers, and/or construction firms.
- Have expertise in designing and implementing simple, scalable processes that solve complex problems.
- Have experience managing complicated dependencies such as logistics and/or supply chains.
- Are relentlessly resourceful and thrive in ambiguous, fast-paced environments.
- Are interested in and thoughtful about the impacts of AGI.
Skills
Gpu Clusters, Hardware, Networking, Capacity Planning, Roadmaps, Risk Management, Automation, Kernels, Scheduling, Supply Chain
Similar jobs
Technical Program Management jobsLeads cross-functional programs that translate model and product needs into scalable, reliable data platform, database, and storage infrastructure. The role requires deep cloud storage and data-stack expertise, architectural judgment, and experience driving complex production programs through adoption.
Leads cross-functional programs that improve developer velocity and deployment excellence across ChatGPT, including CI, testing, release infrastructure, progressive rollouts, and operational validation. Requires strong technical program leadership, systems thinking, and data-driven execution across engineering teams.
Leads cross-functional programs for Chat capacity planning and model deployment, connecting demand forecasting, serving capacity, launch readiness, rollout coordination, and post-deployment learning. Requires technical program leadership in infrastructure, distributed systems, model serving, or large-scale deployment environments.
Leads technical strategy and cross-functional execution for enterprise products and AI workflows, influencing architecture, launch readiness, security, governance, and customer adoption. The role requires strong technical fluency, enterprise software experience, product judgment, and influence across engineering and business functions.
Leads high-stakes technical programs that translate AI safety and product priorities into safeguards, deployment readiness, and measurable outcomes across model, infrastructure, API, and product environments. Requires strong technical judgment, program execution, product judgment, and cross-functional influence.