Technical Account Manager (TAM), AI Factory
Serves as primary technical owner for strategic enterprise AI infrastructure customers, managing compute, networking, storage, and facilities for large-scale GPU deployments. Requires 5+ years customer-facing experience with deep GPU/HPC expertise and observability tools.
About the job
Responsibilities
- Serve as the named technical point of contact for a dedicated strategic customer, owning the end-to-end technical relationship across compute, networking, storage, and facilities
- Drive structured engagement through regular cadences including status reporting, technical steering meetings, and executive business reviews
- Translate customer operational feedback into actionable input for Engineering, Product, and Infrastructure roadmaps
- Lead issue lifecycle management, escalation, and RCA authorship across all infrastructure domains in partnership with Support, SRE, DC Ops, and Engineering teams
- Own end-to-end RMA coordination and hardware lifecycle management, including acceptance testing, spare inventory management, and hardware health reporting for large-scale GPU deployments
- Maintain deep technical expertise across the customer's infrastructure stack — GPU compute, high-speed fabric, and large-scale storage systems — advising on configuration, operational best practices, and incident resolution
- Own the observability strategy for the customer estate, including alert policy definition, dashboard development, and proactive health management across all infrastructure layers
- Coordinate DC operations and facilities events in partnership with internal teams and hosting providers, ensuring SLA compliance and cluster availability
- Act as project manager for all capacity expansions, owning the full node deployment lifecycle from freight receipt through production acceptance
Qualifications
- 5+ years in a customer-facing technical role, with 2+ years in dedicated technical account management or solutions architecture for large-scale AI or HPC infrastructure
- Deep expertise in GPU infrastructure — GPU health diagnostics, RMA workflows, and hardware acceptance testing
- Hands-on experience with large-scale Ethernet and InfiniBand fabric architecture
- Working knowledge of enterprise storage systems, including high-density NVMe, parallel file systems, and metadata infrastructure
- Experience with DC operations, facilities coordination, and hosting provider SLA management
- Strong ownership mindset for incident management, RCA authorship, and executive-level customer communication
- Proficiency in infrastructure monitoring and observability tooling (Prometheus, Grafana, or equivalent)
- Proven ability to manage multiple concurrent workstreams with hyperscaler-level rigor and communication standards
- Proficiency in Python, Bash, or infrastructure automation tools preferred
Compensation
US base salary range: $260-290K OTE + equity + benefits
Skills
Gpu Infrastructure, Ethernet, InfiniBand, Nvme, Parallel File Systems, Prometheus, Grafana, Python, Bash, Rca, Rma, SRE, Dc Operations
Similar jobs
Customer Success jobsOwn partner onboarding, adoption, support, feedback, and operational processes for Suno’s beta partnerships. The role requires 4+ years in customer success, partner operations, technical account management, or program management, plus technical fluency and strong organizational skills.
Own enterprise client relationships across Managed Services, coordinating cross-functional delivery, timelines, reporting, risk management, retention, and expansion. The role requires experience leading complex engagements, advising executive stakeholders, and translating SEO/AEO work into measurable business outcomes.
Supports enterprise customers deploying and scaling AI inference workloads on a cloud platform, providing technical guidance, customer advocacy, training, and incident support. Requires a bachelor’s degree and 2+ years supporting cloud, AI/ML, developer platform, or infrastructure customers.
The ISV Success Manager nurtures software vendor relationships and guides partners through technical integration with Okta and Auth0. The role requires experience in technical partner or customer-facing functions, strong identity and security knowledge, and the ability to influence product, engineering, and business stakeholders.
Owns the customer relationship for a major government account, driving adoption, retention, expansion, and executive engagement while leading Partner Engagement Managers. Requires staff-planning experience, strong technical aptitude, government stakeholder management, and a current Top Secret clearance with SCI eligibility.