# Senior HPC Storage Engineer

**Company:** [Runpod](https://hotfix.jobs/companies/runpod)
**Location:** Remote
**Role:** DevOps / SRE
**Salary:** $180k – $260k/yr
**Experience:** 8+ years
**Skills:** Linux, Ceph, Minio, Lustre, Nvme, Nfs, Iscsi, Nvme-Of, Go, Python, Rust, Kubernetes, Csi Drivers, Prometheus, Grafana
**Posted:** 2026-09-09

> Own the design, scaling, reliability, and automation of a multi-region storage platform supporting AI workloads. The role requires 8+ years of production infrastructure or storage engineering experience, distributed storage expertise, strong Linux and networking knowledge, and production programming skills.

## Job Description

## Responsibilities
- Own capacity, durability, availability, and performance for network volumes, local NVMe, and S3-compatible object storage.
- Tune the full I/O path, including device and filesystem configuration, caching, read-ahead, replication, erasure coding, and client mount behavior.
- Diagnose production performance issues end to end.
- Lead capacity expansions, hardware refreshes, migrations, and rebalances without customer-visible disruption.
- Design and tune storage network paths, including high-throughput east-west fabrics, MTU and jumbo frames, congestion and flow control, multipath, and NIC/offload configuration.
- Optimize RDMA/RoCE and high-speed InfiniBand/Ethernet fabrics for storage traffic.
- Write production code in Go, Python, Rust, or similar for control-plane services, provisioning, data movement, and monitoring.
- Build and extend control-plane, S3-compatible, CSI, Kubernetes, vendor, and cloud-provider APIs.
- Automate operational workflows and manage infrastructure as code through code review, testing, and CI.
- Instrument storage fleets and build metrics, dashboards, SLOs, and alerts for IOPS, throughput, latency, errors, retries, capacity, and tenant consumption.
- Participate in storage-system on-call rotations and lead blameless post-incident follow-up.

## Requirements
- 8+ years of experience in infrastructure, storage, or systems engineering, with substantial ownership of production storage at scale.
- Deep practical experience with at least one distributed storage system, such as Ceph, MinIO, Lustre, GPFS/Spectrum Scale, MooseFS, WekaFS, VAST, or ZFS-based systems.
- Strong Linux internals and storage-stack knowledge, including block devices, filesystems, NVMe, page cache, I/O schedulers, NFS/SMB, and iSCSI/NVMe-oF.
- Experience building or operating S3-compatible object storage services.
- Strong networking fundamentals and experience tuning networks for storage workloads.
- Proficiency shipping production code in Go, Python, Rust, or similar languages.
- Hands-on observability experience with Prometheus, Grafana, Datadog, or equivalent, including metric design.
- Track record of performance analysis and debugging under production pressure.
- Ability to work independently, take ownership across team boundaries, continuously improve systems, and collaborate effectively.

## Nice-to-haves
- Storage experience for AI/ML workloads, including checkpointing, dataset streaming, model-weight distribution, GPU-adjacent data locality, or GPUDirect Storage.
- Kubernetes storage internals, including CSI drivers, PV/PVC lifecycles, StatefulSets, and local persistent volumes.
- Bare-metal and colocation experience, including hardware selection, vendor management, firmware, and physical failure domains.
- Experience operating multi-tenant systems with isolation, fairness, and QoS requirements.
- Experience at a fast-growing cloud or infrastructure provider.

## Compensation and Benefits
- Base salary range: **$180,000–$260,000 USD annually**.
- Equity through stock options.
- Medical, dental, and vision plans.
- Flexible paid time off.
- $1,200 home office and equipment stipend.

## Similar jobs

- [Senior Site Reliability Engineer - Linux Systems & Application Observability](https://hotfix.jobs/jobs/f3e1c6e1-7d10-4e98-9005-e0cd58837354) - tastytrade - Chicago, IL - $180k – $200k/yr
- [Senior Platform Engineer](https://hotfix.jobs/jobs/396237d0-dce2-433a-ba7b-1852c20a74eb) - Sprig - San Francisco, CA - $180k – $260k/yr
- [Senior Platform Software Engineer](https://hotfix.jobs/jobs/c1a5fdfa-4570-43d0-9a7f-d777288315e3) - Camber - New York, NY - $180k – $230k/yr
- [Senior Site Reliability Engineer, Colorado Springs](https://hotfix.jobs/jobs/c3d2b727-7792-4fcd-bb4c-3c6d164c72e6) - Onebrief - Colorado Springs, CO - $180k – $220k/yr
- [Senior Infrastructure Software Engineer](https://hotfix.jobs/jobs/a690500a-c87a-4f20-850d-dc77900b33c3) - Lightning AI - New York, NY - $180k – $220k/yr

**Apply:** https://hotfix.jobs/jobs/72f52714-20f3-4c62-b022-2468417a44d6
**Canonical:** https://hotfix.jobs/jobs/72f52714-20f3-4c62-b022-2468417a44d6