Token-as-a-Service Technical Program Manager
Leads end-to-end delivery of external compute capacity into production-ready tokens for OpenAI model workloads. Drives cross-functional programs across engineering, partners, and operations, requiring 8+ years TPM experience and strong infrastructure knowledge.
About the job
Key Responsibilities
- Lead end-to-end delivery programs that convert external infrastructure capacity into production-ready token supply.
- Own readiness across compute, storage, networking, security, and operational dependencies for third-party environments.
- Build integrated plans across internal engineering teams and external partners with clear milestones, owners, risks, and critical paths.
- Drive launch execution for new partner regions, clusters, and capacity expansions.
- Create operating mechanisms that measure deployed capacity versus usable token output.
- Identify bottlenecks preventing token generation (network constraints, hardware readiness, software enablement, partner delays, etc.) and drive resolution.
- Coordinate with capacity planning and finance teams to prioritize the highest ROI capacity opportunities.
- Establish executive-level reporting on delivery status, risks, and token ramp forecasts.
- Improve repeatability of partner onboarding, technical integration, and scaling motions.
- Manage escalations across internal and external stakeholders during high-severity delivery issues.
- Translate ambiguous infrastructure constraints into clear execution plans.
- Help define the long-term operating model for Token-as-a-Service across Stargate and 3P ecosystems.
Qualifications
- 8+ years of Technical Program Management, Engineering Program Management, or Infrastructure Delivery experience.
- Experience leading large-scale technical programs involving cloud, data center, networking, hardware, or distributed systems.
- Strong understanding of compute infrastructure, clusters, networking, storage, and production systems.
- Proven ability to drive cross-functional execution across engineering, operations, finance, and external vendors.
- Experience managing executive stakeholders and communicating complex tradeoffs clearly.
- Strong analytical skills with ability to reason about utilization, throughput, capacity, and operational metrics.
- Comfortable operating in ambiguous, fast-scaling environments.
- Strong written and verbal communication skills.
- High ownership mentality with bias toward action.
- Experience working with external providers, strategic partners, or hyperscalers is highly preferred.
Preferred Skills
- Experience with GPU clusters, AI infrastructure, or large-scale model serving environments.
- Familiarity with token economics, inference capacity planning, or workload scheduling.
- Experience scaling global infrastructure through third-party providers.
- Background in systems engineering, networking, or hardware deployment programs.
- Experience building new operational models in high-growth environments.
Skills
Cloud Infrastructure, Data Centers, Networking, Hardware, Distributed Systems, Gpu Clusters, AI Infrastructure, Compute Infrastructure, Storage Systems, Capacity Planning
Similar jobs
Technical Program Management jobsOwns portfolio-level milestone tracking, reporting, data quality, and capacity projections for large-scale data center delivery programs. The role requires 7+ years in project controls or scheduling, strong scheduling and BI-tool fluency, and the ability to communicate insights to executive and technical audiences.
Leads cross-functional enterprise-readiness programs spanning security, compliance, identity, data governance, and spend controls across product surfaces. Requires 8+ years of technical program management experience and strong experience with regulated customers and senior stakeholder alignment.
Leads cross-functional programs for billing platform foundations, commercial launches, promotions, charge-pipeline changes, and payments operations. Requires 6+ years of technical program management experience and the ability to coordinate engineering, finance, treasury, product, and support teams.
Leads and scales the security incident management lifecycle for Detection & Response, serving as incident commander and driving post-incident accountability, trend analysis, systemic improvements, and cross-functional coordination. Requires 7+ years of relevant experience, strong analytical and communication skills, and a bachelor’s degree or equivalent experience.
Leads end-to-end delivery of custom AI-agent deployments for enterprise customers in regulated industries, coordinating technical teams, executive stakeholders, product scoping, architecture decisions, and value measurement. Requires production AI/ML deployment experience, enterprise delivery expertise, and strong executive presence.