HPC Infrastructure Engineer - GPU Clusters
Operates and improves large-scale NVIDIA GPU infrastructure supporting AI model training, spanning provisioning, scheduling, networking, storage, automation, performance tuning, and hardware operations. Requires production Linux or GPU environment experience and strong systems and infrastructure skills.
About the job
Responsibilities
- Operate and improve NVIDIA GPU clusters across bare-metal and rented capacity, including provisioning, scheduling, monitoring, upgrades, and capacity planning.
- Build automation for node health checks, automated draining and remediation, and burn-in pipelines for new capacity.
- Own the infrastructure stack beneath training code, including OS images, NVIDIA drivers, CUDA, container runtimes, NCCL, and InfiniBand/RoCE networking.
- Run and tune Slurm or similar job scheduling systems for fair, high-throughput compute access.
- Build and maintain high-performance storage for datasets and checkpoints.
- Diagnose and resolve performance issues involving stragglers, degraded links, thermal conditions, and unreliable GPUs.
- Evaluate rented GPU capacity, benchmark providers, validate performance, and enforce SLAs.
- Perform hands-on hardware work, including racking, cabling, diagnostics, and coordination with datacenter staff and vendors.
- Maintain secure-by-default clusters through access control, network isolation, and secrets management.
Requirements
- Production experience operating large-scale Linux server or GPU environments.
- Strong knowledge of the NVIDIA stack, including drivers, CUDA, NCCL, and DCGM, or deep systems experience with the ability to learn hardware stacks quickly.
- Experience with bare-metal environments, server hardware, and high-speed networking.
- Automation development skills in Python and/or Bash.
- Experience with infrastructure-as-code tools such as Ansible or Terraform.
- Ability to investigate metrics, logs, and PromQL data to diagnose infrastructure problems.
- Willingness to own work end to end and perform datacenter-related tasks when needed.
Nice to Have
- Infrastructure experience supporting ML training workloads, including distributed-training failure modes and checkpointing patterns.
- Experience evaluating and working with GPU cloud providers.
- Experience with parallel filesystems such as WEKA or VAST, or large-scale object storage.
- Experience with BMC/IPMI/Redfish automation and PXE provisioning at scale.
- Awareness of power and cooling considerations for dense GPU deployments.
Skills
Linux, Nvidia Gpu, CUDA, Nccl, Dcgm, InfiniBand, Roce, Slurm, Python, Bash, Ansible, Terraform, Promql, Pxe, Redfish
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.