# HPC Infrastructure Engineer - GPU Clusters

**Company:** [ElevenLabs](https://hotfix.jobs/companies/elevenlabs)
**Location:** Remote
**Role:** DevOps / SRE
**Skills:** Linux, Nvidia Gpu, CUDA, Nccl, Dcgm, InfiniBand, Roce, Slurm, Python, Bash, Ansible, Terraform, Promql, Pxe, Redfish
**Posted:** 2026-09-03

> Operates and improves large-scale NVIDIA GPU infrastructure supporting AI model training, spanning provisioning, scheduling, networking, storage, automation, performance tuning, and hardware operations. Requires production Linux or GPU environment experience and strong systems and infrastructure skills.

## Job Description

## Responsibilities
- Operate and improve NVIDIA GPU clusters across bare-metal and rented capacity, including provisioning, scheduling, monitoring, upgrades, and capacity planning.
- Build automation for node health checks, automated draining and remediation, and burn-in pipelines for new capacity.
- Own the infrastructure stack beneath training code, including OS images, NVIDIA drivers, CUDA, container runtimes, NCCL, and InfiniBand/RoCE networking.
- Run and tune Slurm or similar job scheduling systems for fair, high-throughput compute access.
- Build and maintain high-performance storage for datasets and checkpoints.
- Diagnose and resolve performance issues involving stragglers, degraded links, thermal conditions, and unreliable GPUs.
- Evaluate rented GPU capacity, benchmark providers, validate performance, and enforce SLAs.
- Perform hands-on hardware work, including racking, cabling, diagnostics, and coordination with datacenter staff and vendors.
- Maintain secure-by-default clusters through access control, network isolation, and secrets management.

## Requirements
- Production experience operating large-scale Linux server or GPU environments.
- Strong knowledge of the NVIDIA stack, including drivers, CUDA, NCCL, and DCGM, or deep systems experience with the ability to learn hardware stacks quickly.
- Experience with bare-metal environments, server hardware, and high-speed networking.
- Automation development skills in Python and/or Bash.
- Experience with infrastructure-as-code tools such as Ansible or Terraform.
- Ability to investigate metrics, logs, and PromQL data to diagnose infrastructure problems.
- Willingness to own work end to end and perform datacenter-related tasks when needed.

## Nice to Have
- Infrastructure experience supporting ML training workloads, including distributed-training failure modes and checkpointing patterns.
- Experience evaluating and working with GPU cloud providers.
- Experience with parallel filesystems such as WEKA or VAST, or large-scale object storage.
- Experience with BMC/IPMI/Redfish automation and PXE provisioning at scale.
- Awareness of power and cooling considerations for dense GPU deployments.

## Similar jobs

- [Software Engineer, Infrastructure](https://hotfix.jobs/jobs/585da47a-e02d-4c63-9816-248a2faa9b5b) - Granica - Remote
- [Capacity Ops Engineer](https://hotfix.jobs/jobs/f1904714-7dd3-4ee3-9e7a-e4fcf52083bd) - Baseten - San Francisco, CA - $170k – $230k/yr
- [IT Security and Automation Engineer](https://hotfix.jobs/jobs/604b87b5-13a2-4bba-88b2-f7d0fbbad141) - Teleport - Remote - $149k – $258k/yr
- [Electrical Field Engineer - Data Center](https://hotfix.jobs/jobs/6bfa0e4c-9ccf-438a-b65e-cd4c6297762c) - Crusoe - Remote - $196k – $235k/yr
- [Software Engineer, Cloud Infrastructure](https://hotfix.jobs/jobs/949677d6-6d57-49e8-acf8-017a14790019) - Beacon AI - San Carlos, CA - $135k – $260k/yr

**Apply:** https://hotfix.jobs/jobs/97795223-2089-4e62-a318-f7197d6dc31b
**Canonical:** https://hotfix.jobs/jobs/97795223-2089-4e62-a318-f7197d6dc31b