Member of Technical Staff - GPU Infrastructure
Designs, deploys, and supports large-scale GPU and HPC infrastructure for customers, including cluster architecture, orchestration, networking, storage, and performance optimization. Requires 3+ years of GPU/HPC experience, production SLURM and Kubernetes expertise, and strong customer-facing technical leadership.
About the job
Responsibilities
- Partner with customers to understand workload requirements and design GPU cluster architectures.
- Create technical proposals and capacity plans for clusters ranging from 100 to 10,000+ GPUs.
- Develop deployment strategies for LLM training, inference, and HPC workloads.
- Present architectural recommendations to technical and executive stakeholders.
- Deploy and configure SLURM and Kubernetes for distributed workloads.
- Implement high-performance networking with InfiniBand, RoCE, and NVLink.
- Optimize GPU utilization, memory management, and inter-node communication.
- Configure Lustre, BeeGFS, and GPFS parallel filesystems for optimal I/O performance.
- Tune system performance from Linux kernel parameters through CUDA configurations.
- Serve as the primary technical escalation point for customer infrastructure issues.
- Diagnose and resolve problems across hardware, drivers, networking, and software.
- Implement monitoring, alerting, and automated remediation systems.
- Provide 24/7 on-call support for critical customer deployments.
- Create runbooks and documentation for customer operations teams.
Requirements
- 3+ years of hands-on experience with GPU clusters and HPC environments.
- Deep production expertise with SLURM and Kubernetes in GPU settings.
- Experience configuring and troubleshooting InfiniBand.
- Strong understanding of NVIDIA GPU architecture, CUDA, and the driver stack.
- Experience with Ansible and Terraform.
- Proficiency in Python, Bash, and systems programming.
- Track record of customer-facing technical leadership.
- Experience with NVIDIA driver installation and troubleshooting, including CUDA, Fabric Manager, and DCGM.
- Experience configuring GPU container runtimes such as Docker, Containerd, and Enroot.
- Linux kernel tuning and performance optimization experience.
- Understanding of network topology design for AI workloads.
- Understanding of power and cooling requirements for high-density GPU deployments.
Nice to Have
- Experience with 1,000+ GPU deployments.
- NVIDIA DGX, HGX, or SuperPOD certification.
- Experience with PyTorch FSDP, DeepSpeed, or Megatron-LM.
- ML framework optimization and profiling experience.
- Experience with AMD MI300 or Intel Gaudi accelerators.
- Contributions to open-source HPC or AI infrastructure projects.
Compensation
- Cash compensation range: $150,000–$300,000, plus equity incentives.
Skills
Gpu Clusters, Hpc, Slurm, Kubernetes, InfiniBand, Roce, Nvlink, CUDA, Nvidia Gpus, Ansible, Terraform, Python, Bash, Linux, Docker
Similar jobs
Solutions Architecture jobsDevelop Redis-powered solutions for enterprise customers, lead technical evaluations and pilots, and partner with Sales on account strategy. Requires 3+ years of pre-sales engineering experience, server-side development and Linux expertise, distributed-systems experience, and strong communication skills.
Leads on-site deployments for hospital staffing software, mapping workflows, driving adoption across clinical and ops teams, and building scalable rollout playbooks to activate revenue and ROI.
Advises enterprise customers on AI adoption by identifying use cases, building prototypes, proving business value, and guiding deployment strategy. Requires hands-on customer advisory experience, strong executive communication, enterprise integration knowledge, and solid AI/LLM expertise.
Build and deliver production AI agent systems for enterprise customers through architecture advising, co-development, and embedded engineering engagements. Requires 4+ years of software engineering experience, deep Python expertise, and 2+ years shipping production agent systems.
Build and operate integrations, custom workflows, and full-stack extensions for large school districts, owning deployments from discovery through production stabilization. The role requires 4+ years of software engineering experience, strong debugging and systems skills, and a bachelor’s degree in a rigorous technical field.