# Technical Program Manager, Compute Qualification

**Company:** [Together AI](https://hotfix.jobs/companies/together-ai)
**Location:** Remote
**Role:** Technical Program Management
**Salary:** $200k – $250k/yr
**Experience:** 5+ years
**Skills:** Technical Program Management, infrastructure program management, data center infrastructure, gpu hardware, high-performance networking, InfiniBand, ethernet, storage systems, power and cooling, Python, SQL, Stakeholder Management, hpc infrastructure, ai training infrastructure, cluster design
**Posted:** 2026-07-20

> Own and run the end-to-end qualification process for new GPU compute capacity at Together AI. Coordinate cross-functional engineering validation of providers across hardware, networking, storage, power/cooling; perform first-pass data analysis and deliver go/no-go recommendations to leadership.

## Job Description

## Responsibilities
- Own and continuously improve the end-to-end qualification process for new compute capacity, from initial provider intake through final go/no-go recommendation.
- Run multiple provider evaluations in parallel, setting timelines, tracking status, and keeping every stakeholder aligned on what is needed and by when.
- Partner with infrastructure engineering, network engineering, data center engineering, and SRE teams to plan and coordinate technical validation, then translate their findings into clear decisions for leadership.
- Review provider technical specifications and questionnaire responses for completeness and accuracy, flagging gaps, inconsistencies, and risks that warrant follow-up.
- Conduct first-pass analysis of provider data yourself: compare specifications across suppliers, sanity-check performance claims, and surface issues before deeper engineering review.
- Maintain the standards, templates, and documentation that define what meets spec across compute, networking, storage, power, cooling, and operational support.
- Build a structured, auditable record of evaluation outcomes that informs sourcing decisions and scales the qualification function as the team grows.
- Deep-dive into critical hardware performance metrics, proactively identifying potential bottlenecks in cluster architecture before they impact our end customers training or inference workloads.
- Conduct diligence and work with engineering teams to make assessments regarding technical and operational resilience.

## Requirements
- 5+ years in technical program or project management, infrastructure program management, or a comparable technical operations role, ideally involving hardware, data center, or large-scale compute environments.
- Proven ability to run multiple complex, cross-functional workstreams to deadline, with strong organization and stakeholder management.
- Working technical fluency across data center infrastructure: server and GPU hardware, high-performance networking (InfiniBand or Ethernet fabrics), storage, and power and cooling fundamentals; enough depth to read a detailed technical specification and know what to question.
- Hands-on comfort with data: able to write scripts or queries (for example, Python or SQL) to compare, validate, and analyze provider specifications and test results independently.
- Excellent written and verbal communication; able to turn dense technical detail into clear recommendations for both engineers and executives.
- Willingness to travel to provider and data center sites as needed.

## Nice to Have
- Experience qualifying, commissioning, or accepting GPU clusters or HPC infrastructure against defined performance and reliability standards.
- Familiarity with AI training and inference infrastructure, including interconnect topologies, cluster bring-up, and acceptance testing.
- Experience in AI/HPC cluster design.
- Background working directly with hardware vendors, colocation providers, or cloud capacity providers.

## Similar roles

- [Technical Program Manager, Enablement](https://hotfix.jobs/jobs/9d58b6ab-8d40-471a-97c2-3725be2cd7fa) - Fluidstack - Austin, TX - $200k – $300k/yr
- [Technical Program Manager, Supply Chain](https://hotfix.jobs/jobs/1cdcad18-a071-4161-a0bd-0999856176ff) - Fluidstack - Austin, TX - $200k – $275k/yr
- [Technical Product Manager, Product Experience](https://hotfix.jobs/jobs/e656bea1-9792-457c-937d-845b307922d9) - character.ai - Redwood City, CA - $200k – $300k/yr
- [Technical Program Manager, Data Center Operations](https://hotfix.jobs/jobs/92f950c2-fed6-4a1e-9b2c-38908f156bbc) - Fluidstack - Austin, TX - $200k – $270k/yr
- [Technical Program Manager, Data](https://hotfix.jobs/jobs/6495bdb7-2735-4b8e-acf4-6d515dd989a1) - Sesame - San Francisco, CA - $200k – $260k/yr

**Apply:** https://hotfix.jobs/jobs/6fcd8fcb-d3ad-44db-9ae2-e4afdc53ca07
**Canonical:** https://hotfix.jobs/jobs/6fcd8fcb-d3ad-44db-9ae2-e4afdc53ca07