Hardware Technical Program Manager, Infrastructure Partner Operations
Lead operational delivery and governance for OpenAI's third-party infrastructure partners (cloud providers and compute vendors). Drive SLAs, metrics, escalations, dashboards, and cross-functional programs to ensure reliability for large-scale AI systems. Requires 7+ years in TPM or infrastructure operations.
226k – 285k/yr
On-site7+ YOETechnical Program Management
About the role
Key Responsibilities
Own operational engagement with third-party infrastructure providers, ensuring consistent execution against operational commitments, SLAs, and performance expectations.
Develop operational governance frameworks with strategic partners, including business reviews, operational scorecards, escalation processes, executive reporting, and performance improvement plans.
Define, track, and continuously improve key operational metrics related to infrastructure availability, deployment execution, incident response, operational health, service quality, and partner performance.
Build dashboards and reporting mechanisms that provide clear visibility into partner operational performance, risks, trends, and areas requiring executive attention.
Drive cross-functional coordination between OpenAI teams and external infrastructure providers to resolve operational issues, remove execution blockers, and improve delivery outcomes.
Lead operational escalations involving infrastructure availability, deployment execution, hardware operations, capacity delivery, or service performance, ensuring timely resolution and clear executive communication.
Establish repeatable operating rhythms with external partners, including weekly operational reviews, executive business reviews, service reviews, action tracking, and long-term improvement initiatives.
Partner with Capacity Planning, Hardware Operations, Networking, Deployment, Reliability Engineering, and Supply Chain teams to ensure external infrastructure providers remain aligned with OpenAI’s operational priorities.
Identify systemic operational risks across partner organizations and proactively drive corrective actions that improve long-term operational effectiveness.
Qualifications
7+ years of experience in Technical Program Management, Infrastructure Operations, Cloud Operations, Service Delivery, or Technical Account Management within large-scale infrastructure environments.
Experience managing operational relationships with external infrastructure providers, cloud service providers, hardware vendors, or strategic technology partners.
Strong understanding of hyperscale cloud infrastructure, data center operations, infrastructure delivery, or large-scale distributed systems.
Experience developing operational KPIs, SLAs, service health metrics, dashboards, and executive reporting for complex technical organizations.
Demonstrated success leading cross-functional operational programs involving both internal stakeholders and external partners.
Strong program management skills with the ability to drive accountability across organizations without direct authority.
Excellent written and verbal communication skills with experience presenting operational performance to senior technical and executive leadership.
Bachelor's degree in Engineering, Computer Science, Information Systems, Operations, or equivalent practical experience.
Preferred Skills
Experience managing cloud infrastructure operations within organizations such as Microsoft Azure, Amazon Web Services (AWS), Google Cloud Platform (GCP), Oracle Cloud Infrastructure (OCI), or other hyperscale cloud providers.
Experience leading operational governance, service delivery, customer success engineering, technical account management, or infrastructure operations for enterprise cloud customers.
Strong understanding of service-level agreements (SLAs), operational KPIs, incident management, escalation processes, root cause analysis, and continuous service improvement methodologies.
Experience building executive dashboards, operational scorecards, business review frameworks, and data-driven performance reporting.
Familiarity with infrastructure operations supporting GPU infrastructure, AI infrastructure, high-performance computing (HPC), or hyperscale data center environments.
Experience managing complex cross-company technical relationships while balancing customer priorities, engineering constraints, and operational execution.
Proven ability to influence senior stakeholders across both internal teams and external partner organizations without direct authority.
Experience driving continuous operational improvements through metrics, process optimization, and structured governance.
Skills
Technical Program Managementinfrastructure operationscloud operationsservice deliverytechnical account managementslasKPIsDashboardsexecutive reportingIncident ManagementRoot Cause AnalysisAWSGCPAzurehpc
Supply Chain Program Manager (SCPM) - AI Infrastructure
OpenAISan Francisco, CA
Owns material readiness and supply chain execution for AI hardware programs including custom silicon, systems, and infrastructure. Drives cross-functional alignment with engineering, sourcing, and suppliers to manage constrained commodities and NPI timelines. Requires 8+ years in hardware supply chain roles.
226k – 285k/yr
Hybrid8+ YOETechnical Program Management
Sr. Technical Program Manager (TPM)
Together AISan Francisco, CA
Leads development and scaling of global GPU infrastructure for AI, owning product roadmaps for observability, storage, networking, and security. Requires 5+ years in AI/ML infrastructure, cloud platforms, and cross-functional leadership in a fast-paced startup.
225k – 265k/yr
Hybrid5+ YOETechnical Program Management
Technical Program Manager, QA Program
FluidstackAustin, TX
Build and run the QA program for gigawatt-scale data center construction, defining enforceable quality gates, managing nonconformances, and delivering actionable quality metrics from factory to energization. Requires experience running QA programs in construction or manufacturing at scale.
222k – 295k/yr
On-site7+ YOETechnical Program Management
Infrastructure Delivery Lead
FluidstackNew York, NY
Own end-to-end delivery of AI compute sites from construction handover to customer-ready GPUs, running integrated schedules, chairing daily delivery rhythms, and owning readiness gates across converging workstreams at hundreds of megawatts per site.
222k – 307k/yr
On-site7+ YOETechnical Program Management
Technical Program Manager Lead, Deployments
FluidstackAustin, TX +3
Lead TPM responsible for running multi-site data center deployment programs at massive scale, owning schedules, cross-site sequencing decisions, accurate leadership reporting, and building the TPM team to deliver GWs of AI compute infrastructure rapidly.