# Member of Technical Staff - GPU Infrastructure

**Company:** [Prime Intellect](https://hotfix.jobs/companies/prime-intellect)
**Location:** San Francisco, CA
**Role:** Solutions Architecture
**Salary:** $150k – $300k/yr
**Experience:** 3+ years
**Skills:** Gpu Clusters, Hpc, Slurm, Kubernetes, InfiniBand, Roce, Nvlink, CUDA, Nvidia Gpus, Ansible, Terraform, Python, Bash, Linux, Docker
**Posted:** 2026-07-08

> Designs, deploys, and supports large-scale GPU and HPC infrastructure for customers, including cluster architecture, orchestration, networking, storage, and performance optimization. Requires 3+ years of GPU/HPC experience, production SLURM and Kubernetes expertise, and strong customer-facing technical leadership.

## Job Description

## Responsibilities
- Partner with customers to understand workload requirements and design GPU cluster architectures.
- Create technical proposals and capacity plans for clusters ranging from 100 to 10,000+ GPUs.
- Develop deployment strategies for LLM training, inference, and HPC workloads.
- Present architectural recommendations to technical and executive stakeholders.
- Deploy and configure SLURM and Kubernetes for distributed workloads.
- Implement high-performance networking with InfiniBand, RoCE, and NVLink.
- Optimize GPU utilization, memory management, and inter-node communication.
- Configure Lustre, BeeGFS, and GPFS parallel filesystems for optimal I/O performance.
- Tune system performance from Linux kernel parameters through CUDA configurations.
- Serve as the primary technical escalation point for customer infrastructure issues.
- Diagnose and resolve problems across hardware, drivers, networking, and software.
- Implement monitoring, alerting, and automated remediation systems.
- Provide 24/7 on-call support for critical customer deployments.
- Create runbooks and documentation for customer operations teams.

## Requirements
- 3+ years of hands-on experience with GPU clusters and HPC environments.
- Deep production expertise with SLURM and Kubernetes in GPU settings.
- Experience configuring and troubleshooting InfiniBand.
- Strong understanding of NVIDIA GPU architecture, CUDA, and the driver stack.
- Experience with Ansible and Terraform.
- Proficiency in Python, Bash, and systems programming.
- Track record of customer-facing technical leadership.
- Experience with NVIDIA driver installation and troubleshooting, including CUDA, Fabric Manager, and DCGM.
- Experience configuring GPU container runtimes such as Docker, Containerd, and Enroot.
- Linux kernel tuning and performance optimization experience.
- Understanding of network topology design for AI workloads.
- Understanding of power and cooling requirements for high-density GPU deployments.

## Nice to Have
- Experience with 1,000+ GPU deployments.
- NVIDIA DGX, HGX, or SuperPOD certification.
- Experience with PyTorch FSDP, DeepSpeed, or Megatron-LM.
- ML framework optimization and profiling experience.
- Experience with AMD MI300 or Intel Gaudi accelerators.
- Contributions to open-source HPC or AI infrastructure projects.

## Compensation
- Cash compensation range: **$150,000–$300,000**, plus equity incentives.

## Similar jobs

- [Solution Architect](https://hotfix.jobs/jobs/940e71b7-3254-46a4-9edb-42cdf3e49102) - Redis - Remote - $150k – $188k/yr
- [Deployment Strategist](https://hotfix.jobs/jobs/71728831-193e-4f41-b6f3-8bc9efd0b7a7) - Vitalize - San Francisco, CA - $150k – $240k/yr
- [Forward Deployed Engineer](https://hotfix.jobs/jobs/ccdcc67c-00c7-47e5-ab90-5303ae6c4e67) - SnapLogic - $150k – $200k/yr
- [Deployed Engineer, Professional Services](https://hotfix.jobs/jobs/942cb249-1514-4a85-83f0-b97e31d7b383) - LangChain - San Francisco, CA - $150k – $215k/yr
- [Forward Deployed Engineer](https://hotfix.jobs/jobs/cb975321-8332-49bb-b097-87c76172e1ef) - Edia - Remote - $150k – $200k/yr

**Apply:** https://hotfix.jobs/jobs/eea94630-3337-499d-a681-bd66e11758aa
**Canonical:** https://hotfix.jobs/jobs/eea94630-3337-499d-a681-bd66e11758aa