Skip to content

Technical Program Manager, Compute Qualification

Own and run the end-to-end qualification process for new GPU compute capacity at Together AI. Coordinate cross-functional engineering validation of providers across hardware, networking, storage, power/cooling; perform first-pass data analysis and deliver go/no-go recommendations to leadership.

About the job

Responsibilities

  • Own and continuously improve the end-to-end qualification process for new compute capacity, from initial provider intake through final go/no-go recommendation.
  • Run multiple provider evaluations in parallel, setting timelines, tracking status, and keeping every stakeholder aligned on what is needed and by when.
  • Partner with infrastructure engineering, network engineering, data center engineering, and SRE teams to plan and coordinate technical validation, then translate their findings into clear decisions for leadership.
  • Review provider technical specifications and questionnaire responses for completeness and accuracy, flagging gaps, inconsistencies, and risks that warrant follow-up.
  • Conduct first-pass analysis of provider data yourself: compare specifications across suppliers, sanity-check performance claims, and surface issues before deeper engineering review.
  • Maintain the standards, templates, and documentation that define what meets spec across compute, networking, storage, power, cooling, and operational support.
  • Build a structured, auditable record of evaluation outcomes that informs sourcing decisions and scales the qualification function as the team grows.
  • Deep-dive into critical hardware performance metrics, proactively identifying potential bottlenecks in cluster architecture before they impact our end customers training or inference workloads.
  • Conduct diligence and work with engineering teams to make assessments regarding technical and operational resilience.

Requirements

  • 5+ years in technical program or project management, infrastructure program management, or a comparable technical operations role, ideally involving hardware, data center, or large-scale compute environments.
  • Proven ability to run multiple complex, cross-functional workstreams to deadline, with strong organization and stakeholder management.
  • Working technical fluency across data center infrastructure: server and GPU hardware, high-performance networking (InfiniBand or Ethernet fabrics), storage, and power and cooling fundamentals; enough depth to read a detailed technical specification and know what to question.
  • Hands-on comfort with data: able to write scripts or queries (for example, Python or SQL) to compare, validate, and analyze provider specifications and test results independently.
  • Excellent written and verbal communication; able to turn dense technical detail into clear recommendations for both engineers and executives.
  • Willingness to travel to provider and data center sites as needed.

Nice to Have

  • Experience qualifying, commissioning, or accepting GPU clusters or HPC infrastructure against defined performance and reliability standards.
  • Familiarity with AI training and inference infrastructure, including interconnect topologies, cluster bring-up, and acceptance testing.
  • Experience in AI/HPC cluster design.
  • Background working directly with hardware vendors, colocation providers, or cloud capacity providers.

Skills

Technical Program Management, Infrastructure Program Management, Data Center Infrastructure, Gpu Hardware, High-Performance Networking, InfiniBand, Ethernet, Storage Systems, Power And Cooling, Python, SQL, Stakeholder Management, Hpc Infrastructure, Ai Training Infrastructure, Cluster Design

Fluidstack

Fluidstack

Austin, TX
Technical Program Manager, Enablement
$200k+/yrOn-site5+ YOETechnical Program Management

Drive adoption of an internal automation platform for AI infrastructure teams by running rollout, training, and enablement programs. Build frontend fixes (React/TS), prototype with AI tools, measure usage, and translate operational friction into shipped improvements.

Scale AI

Scale AI

Columbia, SC
Technical Program Manager , Public Sector
$199k+/yrHybrid5+ YOETechnical Program Management

Leads cross-functional delivery of AI/ML programs for public-sector cyber customers, managing accounts, datasets, deployments, and issue resolution. Requires cybersecurity experience, an active Top Secret clearance with polygraph, technical education, and willingness to work onsite in Columbia four days weekly.

Valon

Valon

New York, NY
Technical Program Manager
$205k+/yrOn-site5+ YOETechnical Program Management

Leads complex, cross-functional technical programs from planning through launch, coordinating engineering and business stakeholders while managing risks, dependencies, and delivery metrics. Requires 5+ years of technical program management experience in complex software development environments.

Anthropic

Anthropic

Boston, MA
Executive Services Program Manager
$195k+/yrHybridTechnical Program Management

Executes personal security and GSIS program workstreams, maintaining policies, reporting, documentation, and stakeholder support under senior direction. The role requires corporate or physical security program experience, strong writing and documentation skills, and a bachelor’s degree or equivalent experience.

OpenAI

OpenAI

San Francisco, CA

Technical Program Manager, Foundations Data and Operations
$207k+/yrHybridTechnical Program Management

Leads complex search research and engineering programs across retrieval, indexing, model training, and infrastructure. The role requires substantial technical experience, strong program execution, product judgment, and the ability to align cross-functional teams.