TPM
Technical Program Manager leading capacity planning, forecasting, allocation, and fleet strategy for Cerebras' high-speed AI inference service. Own rolling capacity models, datacenter bring-up, weekly reviews with SRE/Product/Eng, tool adoption, and strategic initiatives to maximize utilization of the world's largest AI chips.
About the job
What you'll own
- Capacity planning and forecasting: Build and maintain the 6/12/26-week rolling capacity model across every cluster. Translate customer contracts and sales pipeline into capacity requirements. Forecast model replicas, system-hours, and spares. Reconcile against actuals weekly and maintain the source-of-truth document.
- New Datacenter Capacity bring-up: Collaborate with datacenter infrastructure and operations teams to support new datacenter bring-up, ensure production readiness, drive engineering efforts and automation for on-time delivery.
- Allocation and cluster placement: Partner with SRE and product teams to run weekly capacity reviews. Decide model placement, re-balancing, customer tenant assignments, cluster absorption of new launches, and freezes. Run weekly capacity and utilization reports. Drive downstream model deployment tasks with SRE.
- Drive capacity planning tool adoption: Partner with console engineering to drive stakeholder adoption of the in-house capacity planning and allocation tool (user acceptance testing, issue resolution, tracking changes, pilot testing, deployment). Contribute to continuous process improvement and development of internal capacity management tools.
- Incident tracking and postmortems: Proactively identify and mitigate capacity bottlenecks, risks, and dependencies. Drive resolution and postmortems for SLA drops due to capacity misallocations.
Key Responsibilities
- Run weekly capacity planning and daily capacity and deployment tracking with Engineering, Product, and Operations teams.
- Own fleet utilization reporting and forecasting.
- Drive capacity planning for new customer deployments and major model launches.
- Drive continuous improvement and stakeholder adoption of new capacity management platform.
- Drive org-level strategic initiatives related to capacity expansion, improving fleet efficiency, and maximizing utilization of available systems.
- Lead planning around major infrastructure events (new customer commits, new model releases, changes to DC/cluster architecture) that impact capacity and fleet utilization; update plans and forecasts accordingly.
- Maintain Jira EPICs and Confluence pages related to capacity planning, reporting, and change management for execution transparency.
Qualifications
- 5+ years of TPM, technical program management, or product operations experience in cloud infrastructure, large-scale ML serving, or hyperscaler capacity planning.
- Experience leading large cross-functional programs involving Engineering, Product, and Operations.
- Comfort with the inference serving stack: model replicas, batching, prefill/decode, KV cache, accelerator scheduling.
- Strong data fluency: SQL, Grafana, basic Python or Flux to pull your own numbers.
- Track record of running a recurring cross-functional ritual involving senior engineers and leadership.
- Direct experience with AI accelerator fleet operations such as Habana, TPU pods, Inferentia, Trainium.
Skills
Capacity Planning, Fleet Management, SQL, Grafana, Python, Jira, Confluence, Inference Serving, Ml Serving, Ai Accelerators, Cross-Functional Program Management
Similar jobs
Technical Program Management jobsLeads cross-functional delivery of commercial aerospace and software programs from proposal through deployment, coordinating engineering, product, business development, and customer stakeholders. Requires 5+ years of complex program management experience and fluency in Agile and Waterfall delivery.
Leads cross-functional delivery of AI/ML programs for public-sector cyber customers, managing accounts, datasets, deployments, and issue resolution. Requires cybersecurity experience, an active Top Secret clearance with polygraph, technical education, and willingness to work onsite in Columbia four days weekly.
Leads cross-functional delivery of AI/ML solutions for national security customers, combining technical program management, stakeholder leadership, analytics, and customer success. Requires an active TS/SCI clearance, generative AI/ML fluency, and regular work at a cleared Boston-area facility.
Leads cross-functional delivery of AI/ML solutions and agentic workflows for public sector and national security customers. Requires technical program management, stakeholder leadership, generative AI expertise, an active TS/SCI clearance, and customer-facing problem-solving experience.
Leads infrastructure delivery workstreams spanning compute, networking, storage, hardware, platform engineering, and data center operations. Requires 5+ years of technical program management experience, strong cross-functional execution, and familiarity with infrastructure environments.