Skip to content
RunpodRunpod

Senior HPC Storage Engineer

Own the design, scaling, reliability, and automation of a multi-region storage platform supporting AI workloads. The role requires 8+ years of production infrastructure or storage engineering experience, distributed storage expertise, strong Linux and networking knowledge, and production programming skills.

About the job

Responsibilities

  • Own capacity, durability, availability, and performance for network volumes, local NVMe, and S3-compatible object storage.
  • Tune the full I/O path, including device and filesystem configuration, caching, read-ahead, replication, erasure coding, and client mount behavior.
  • Diagnose production performance issues end to end.
  • Lead capacity expansions, hardware refreshes, migrations, and rebalances without customer-visible disruption.
  • Design and tune storage network paths, including high-throughput east-west fabrics, MTU and jumbo frames, congestion and flow control, multipath, and NIC/offload configuration.
  • Optimize RDMA/RoCE and high-speed InfiniBand/Ethernet fabrics for storage traffic.
  • Write production code in Go, Python, Rust, or similar for control-plane services, provisioning, data movement, and monitoring.
  • Build and extend control-plane, S3-compatible, CSI, Kubernetes, vendor, and cloud-provider APIs.
  • Automate operational workflows and manage infrastructure as code through code review, testing, and CI.
  • Instrument storage fleets and build metrics, dashboards, SLOs, and alerts for IOPS, throughput, latency, errors, retries, capacity, and tenant consumption.
  • Participate in storage-system on-call rotations and lead blameless post-incident follow-up.

Requirements

  • 8+ years of experience in infrastructure, storage, or systems engineering, with substantial ownership of production storage at scale.
  • Deep practical experience with at least one distributed storage system, such as Ceph, MinIO, Lustre, GPFS/Spectrum Scale, MooseFS, WekaFS, VAST, or ZFS-based systems.
  • Strong Linux internals and storage-stack knowledge, including block devices, filesystems, NVMe, page cache, I/O schedulers, NFS/SMB, and iSCSI/NVMe-oF.
  • Experience building or operating S3-compatible object storage services.
  • Strong networking fundamentals and experience tuning networks for storage workloads.
  • Proficiency shipping production code in Go, Python, Rust, or similar languages.
  • Hands-on observability experience with Prometheus, Grafana, Datadog, or equivalent, including metric design.
  • Track record of performance analysis and debugging under production pressure.
  • Ability to work independently, take ownership across team boundaries, continuously improve systems, and collaborate effectively.

Nice-to-haves

  • Storage experience for AI/ML workloads, including checkpointing, dataset streaming, model-weight distribution, GPU-adjacent data locality, or GPUDirect Storage.
  • Kubernetes storage internals, including CSI drivers, PV/PVC lifecycles, StatefulSets, and local persistent volumes.
  • Bare-metal and colocation experience, including hardware selection, vendor management, firmware, and physical failure domains.
  • Experience operating multi-tenant systems with isolation, fairness, and QoS requirements.
  • Experience at a fast-growing cloud or infrastructure provider.

Compensation and Benefits

  • Base salary range: $180,000–$260,000 USD annually.
  • Equity through stock options.
  • Medical, dental, and vision plans.
  • Flexible paid time off.
  • $1,200 home office and equipment stipend.

Skills

Linux, Ceph, Minio, Lustre, Nvme, Nfs, Iscsi, Nvme-Of, Go, Python, Rust, Kubernetes, Csi Drivers, Prometheus, Grafana

tastytrade

tastytrade

Chicago, IL

Senior Site Reliability Engineer - Linux Systems & Application Observability
$180k+/yrHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.

Sprig

Sprig

San Francisco, CA

Senior Platform Engineer
$180k+/yrHybrid6+ YOEDevOps / SRE

Own and modernize the build, CI, test automation, and ephemeral environment platform for a large TypeScript, React, and Go monorepo. The role requires 6+ years of large-scale build-system experience, strong Bazel or comparable tooling expertise, and deep knowledge of hermetic, reproducible development workflows.

Camber

Camber

New York, NY

Senior Platform Software Engineer
$180k+/yrOn-site6+ YOEDevOps / SRE

Senior platform engineer responsible for reliable, secure, and scalable infrastructure, developer tooling, observability, and AI enablement. The role requires 6+ years in platform engineering, SRE, or DevOps, with strong AWS and incident leadership experience.

Onebrief

Onebrief

Colorado Springs, CO

Senior Site Reliability Engineer, Colorado Springs
$180k+/yrOn-site5+ YOEDevOps / SRE

Own reliability, scalability, security, observability, and incident response for mission-critical applications across Kubernetes, AWS, and on-premise DoD environments. Requires an active Top Secret clearance and at least five years of infrastructure-focused SRE, DevOps, or platform engineering experience.

Lightning AI

Lightning AI

New York, NY

Senior Infrastructure Software Engineer
$180k+/yrHybrid8+ YOEDevOps / SRE

Build and operate production software, APIs, and automation for large-scale bare-metal and GPU infrastructure. The role requires 8+ years of software or infrastructure engineering experience, strong Python and Linux skills, and expertise in provisioning, lifecycle management, and reliability.