Staff Technical Program Manager, Managed Intelligence
Own end-to-end program delivery for Crusoe's Managed Inference platform, coordinating model engineering, IaaS, and data center ops to deliver reliable LLM inference at scale. Requires 7+ years TPM experience with deep knowledge of LLM serving, model onboarding, and multi-tenant systems.
About the job
What You'll Be Working On
- End-to-end program delivery: Own multi-quarter release planning, dependency governance, and executive communication across the Managed Inference platform.
- Complex, high-risk program management: Drive model version rollouts, inference optimization campaigns, SLA readiness for new GPU hardware, and multi-tenant capacity planning from kickoff through delivery.
- Cross-functional alignment: Coordinate across Model Engineering, IaaS, Cloud Foundations, Data Center Operations, and external model providers to keep programs on track and unblocked.
- Proactive risk identification: Surface risks across model serving, reliability, capacity constraints, and vendor timelines before they become program-level problems.
- Execution frameworks and dashboards: Build lightweight, scalable TPM frameworks suited to Crusoe's pace; maintain real-time execution dashboards and deliver crisp, data-driven executive updates.
- Phase 0 planning for model onboarding: Own pre-launch planning for model onboarding on new GPU generations, including firmware and driver readiness, CUDA and ROCm stack validation, and commissioning criteria for inference workloads.
- Stakeholder leadership: Drive alignment and push back effectively across engineering, product, and operations leadership -- including highly technical stakeholders who have not previously worked with a TPM.
What You'll Bring to the Team
- 7+ years of experience as a Technical Program Manager in fast-paced technical environments, with a track record of owning complex programs end-to-end across engineering and product organizations.
- LLM inference and model serving knowledge: Working familiarity with batching strategies, quantization approaches, and the tradeoffs that govern latency, throughput, and cost at production scale.
- Multi-tenant systems experience: Familiarity with isolation, quota management, and SLA enforcement across concurrent workloads.
- Fine-tuning and alignment awareness: Sufficient familiarity with fine-tuning and alignment workflows to govern program timelines, identify technical risks, and coordinate across the teams that own them.
- Low-structure execution: Proven ability to build execution models in environments where the process did not yet exist, and make them stick with teams that didn't ask for them.
- Executive communication: Exceptional written and verbal communication for delivering clear, data-driven, decision-oriented updates to executive stakeholders.
- AI tool integration: Active, daily use of AI tools to improve program execution, risk detection, and communication.
- Cross-functional influence: Proven ability to drive alignment across engineering, product, and infrastructure leadership without direct authority, including with highly technical stakeholders.
Bonus Points
- 1+ years of experience working with teams building platforms or services for AI inference and/or training.
- Direct experience governing model onboarding programs across GPU generations, including firmware, driver, and stack validation.
- Experience coaching or mentoring junior TPMs in a high-growth technical environment.
- Exposure to multi-site or globally distributed engineering teams.
- Background at a Series D to Series F company or a high-performing team within a hyperscaler focused on AI infrastructure.
Benefits
- Competitive compensation and equity packages
- Restricted Stock Units
- Paid time off, paid holidays & leave of absence programs
- Comprehensive health, dental & vision insurance
- Employer contributions to HSA account
- Paid parental leave
- Paid life insurance, short-term and long-term disability
- Professional development & tuition reimbursement
- Mental health & wellness support
- Commuter benefits (parking & transit)
- Cell phone stipend
- 401(k) Retirement plan with company match up to 4% of salary
- Volunteer time off
- Global travel insurance & emergency assistance
- Daily meals allowance
- Additional perks & programs specific to location
Skills
Technical Program Management, Llm Inference, Model Serving, Batching Strategies, Quantization, Multi-Tenant Systems, Sla Management, CUDA, Rocm, Gpu Infrastructure, AI Tools, Cross-Functional Coordination, Executive Communication, Risk Management
Similar jobs
Technical Program Management jobsOwns the electrical and electrical-distribution program scope for automotive vehicle programs developed with external partners, coordinating harness, E/E integration, validation, release control, milestones, and executive reporting. Requires 10+ years in automotive technical program management or related engineering leadership.
Owns the strategy, governance, architecture, and hands-on development of Fetch’s People data and analytics foundation. The role builds semantic models, reporting products, and operating processes while enabling trusted metrics, self-service analytics, and future AI automation.
Leads high-impact technical programs for the Duolingo English Test, coordinating engineering, product, data, legal, security, and external stakeholders from planning through launch. Requires staff-level autonomy, strong technical communication, and experience managing complex programs in agile software environments.
Leads cross-functional programs delivering end-to-end autonomous mobility features across vehicle, autonomy, and operations teams. Requires 10+ years in engineering or program management, experience with complex product development, and strong stakeholder communication.
Leads complex, cross-organizational AI/ML programs for autonomous mobility, translating strategy into roadmaps, managing risks and metrics, and aligning senior engineering stakeholders. Requires 10–12+ years of engineering program management experience and deep expertise in relevant autonomy or AI/ML domains.