HPC/ GPU Cluster Architect
Designs, architects, and scales production GPU/HPC clusters globally. Debugs hardware/software issues, automates operations, and mentors juniors. Requires 5+ years experience and hybrid SF presence.
About the job
Responsibilities
- Architect and deploy new GPU/HPC clusters around the world
- Keep clusters running smoothly
- Participate in on-call rotation
- Deploy new environments and fix issues
- Lean into automation for deployments at scale
- Mentor junior engineers and shape team culture
Requirements
- 5+ years experience designing, architecting, and scaling HPC or GPU compute clusters in production
- Deep understanding of server hardware: GPUs, NICs, PCIe, memory, thermals, power
- Comfortable debugging performance/reliability across hardware, OS, drivers, networking
- Automate fleet operations (provisioning, monitoring, remediation) with infrastructure-as-code
- Generate strong operational documentation and runbooks
- Open to SF office 3-4 days/week and domestic travel
Nice to Haves
- Data center operations: power, cooling, colo/vendor engagements
- Strong Linux sysadmin: kernel drivers, RDMA tuning, performance analysis
- Schedulers/orchestration: Slurm, Kubernetes
- Virtualization: KVM, QEMU, libvirt
- Telemetry for predictive hardware failure
- High-speed fabrics: InfiniBand, RoCEv2 Ethernet
Skills
GPU, Hpc, Kubernetes, Slurm, InfiniBand, Rdma, Linux, Pcie, Infrastructure As Code, Kvm
Similar jobs
DevOps / SRE jobsBuild and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.
Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.
Owns reliability standards, incident management, observability, failure testing, and automation for a high-throughput AI infrastructure platform. The role requires deep Linux, networking, software, cloud-native, and distributed-systems experience, along with the ability to influence teams across the organization.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.