Head of Global Compute Supply & Platform Strategy
Lead global compute strategy and platform operations for frontier AI/robotics training at Luma. Own multi-year capacity planning, vendor partnerships, megawatt-scale capital deployment, and infrastructure leadership for 10k+ accelerator fleets.
About the job
What You'll Do
- Architect Multi-Year Compute Strategy: Lead capacity planning, global vendor and cloud partnerships, on-prem vs. cloud mix, and accelerator supply chain roadmaps (H/B-series GPUs, custom silicon evaluation).
- Direct the Platform Org: Provide strategic leadership to our infrastructure, distributed systems, and datacenter operations teams—scaling the organization to support next-generation compute demands.
- Maximize Fleet Utilization: Oversee the architectural efficiency of our cluster configurations to deliver >50% Model Flops Utilization (MFU) on flagship training runs.
- Command a Megawatt Budget: Negotiate, secure, and operate our largest-scale capital deployments for compute infrastructure, partnering directly with Finance to optimize unit economics and risk management.
- Unify Global Capacity: Champion the platform strategy that enables world-model training, heavy simulation rollouts, and real-time on-robot inference to seamlessly share a single, elastic fleet.
- Act as Principal Executive Interface: Serve as the primary commercial and strategic bridge to NVIDIA, AMD, hyperscalers, and frontier silicon vendors.
Qualifications
- 10+ years of engineering leadership experience in large-scale distributed systems, infrastructure, or technical supply chain, with a proven track record of leading compute platform strategy at a frontier AI lab, hyperscaler, or major autonomy program.
- Deep technical & commercial fluency in high-performance cluster topology, high-speed interconnects (InfiniBand/RoCE), large-scale data systems, and the economics of distributed training architectures.
- Direct operational oversight of 10k+ accelerator environments in high-performance production settings.
Preferred Qualifications
- Experience orchestrating capital or infrastructure for training runs at the >100B-parameter or >100k-GPU-day scale.
- Familiarity with the unique capacity and latency demands of edge-to-cloud inference and real-time autonomous systems.
Compensation
- The base pay range for this role is $250,000 – $450,000 per year.
Skills
Large-Scale Distributed Systems, Infrastructure, Technical Supply Chain, Compute Platform Strategy, High-Performance Cluster Topology, InfiniBand, Roce, Large-Scale Data Systems, Distributed Training Architectures, Accelerator Environments, Nvidia, Amd, Hyperscalers, Capital Planning, Datacenter Operations
Similar jobs
Engineering Management jobsLeads the organization responsible for deploying, sustaining, supporting, and improving integrated Hivemind software and hardware products in customer environments. Requires 15 years of technical operations or lifecycle leadership experience, systems engineering expertise, and experience building operational organizations.
Leads a 12-person application security practice, combining hands-on code auditing and vulnerability research with client delivery, staffing, profitability, and engineer development. Requires 10+ years in security and proficiency in at least four relevant programming languages.
Leads the engineering organization for a deployed Group 3 VTOL UAS program, setting technical strategy, governing architecture and development practices, and scaling multidisciplinary teams. Requires 15+ years of related experience, aerospace or adjacent technical expertise, and a bachelor's degree or equivalent experience.
Leads the entire engineering organization, owning technical vision, architecture, execution standards, security, budget, and organizational scaling. The role requires extensive software engineering and engineering management experience, cloud-native architecture expertise, and a record of building high-performing teams.
Leads Okta’s security GRC organization, overseeing enterprise cyber risk, AI governance, global compliance, audits, vendor risk, and engineering-driven remediation. The role requires 10+ years of progressive Security GRC leadership, cloud technology experience, AI governance expertise, and a bachelor’s degree or equivalent experience.