Senior HPC Storage Engineer
Own the design, scaling, reliability, and automation of a multi-region storage platform supporting AI workloads. The role requires 8+ years of production infrastructure or storage engineering experience, distributed storage expertise, strong Linux and networking knowledge, and production programming skills.
About the job
Responsibilities
- Own capacity, durability, availability, and performance for network volumes, local NVMe, and S3-compatible object storage.
- Tune the full I/O path, including device and filesystem configuration, caching, read-ahead, replication, erasure coding, and client mount behavior.
- Diagnose production performance issues end to end.
- Lead capacity expansions, hardware refreshes, migrations, and rebalances without customer-visible disruption.
- Design and tune storage network paths, including high-throughput east-west fabrics, MTU and jumbo frames, congestion and flow control, multipath, and NIC/offload configuration.
- Optimize RDMA/RoCE and high-speed InfiniBand/Ethernet fabrics for storage traffic.
- Write production code in Go, Python, Rust, or similar for control-plane services, provisioning, data movement, and monitoring.
- Build and extend control-plane, S3-compatible, CSI, Kubernetes, vendor, and cloud-provider APIs.
- Automate operational workflows and manage infrastructure as code through code review, testing, and CI.
- Instrument storage fleets and build metrics, dashboards, SLOs, and alerts for IOPS, throughput, latency, errors, retries, capacity, and tenant consumption.
- Participate in storage-system on-call rotations and lead blameless post-incident follow-up.
Requirements
- 8+ years of experience in infrastructure, storage, or systems engineering, with substantial ownership of production storage at scale.
- Deep practical experience with at least one distributed storage system, such as Ceph, MinIO, Lustre, GPFS/Spectrum Scale, MooseFS, WekaFS, VAST, or ZFS-based systems.
- Strong Linux internals and storage-stack knowledge, including block devices, filesystems, NVMe, page cache, I/O schedulers, NFS/SMB, and iSCSI/NVMe-oF.
- Experience building or operating S3-compatible object storage services.
- Strong networking fundamentals and experience tuning networks for storage workloads.
- Proficiency shipping production code in Go, Python, Rust, or similar languages.
- Hands-on observability experience with Prometheus, Grafana, Datadog, or equivalent, including metric design.
- Track record of performance analysis and debugging under production pressure.
- Ability to work independently, take ownership across team boundaries, continuously improve systems, and collaborate effectively.
Nice-to-haves
- Storage experience for AI/ML workloads, including checkpointing, dataset streaming, model-weight distribution, GPU-adjacent data locality, or GPUDirect Storage.
- Kubernetes storage internals, including CSI drivers, PV/PVC lifecycles, StatefulSets, and local persistent volumes.
- Bare-metal and colocation experience, including hardware selection, vendor management, firmware, and physical failure domains.
- Experience operating multi-tenant systems with isolation, fairness, and QoS requirements.
- Experience at a fast-growing cloud or infrastructure provider.
Compensation and Benefits
- Base salary range: $180,000–$260,000 USD annually.
- Equity through stock options.
- Medical, dental, and vision plans.
- Flexible paid time off.
- $1,200 home office and equipment stipend.
Skills
Linux, Ceph, Minio, Lustre, Nvme, Nfs, Iscsi, Nvme-Of, Go, Python, Rust, Kubernetes, Csi Drivers, Prometheus, Grafana
Similar jobs
DevOps / SRE jobsSenior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.
Own and modernize the build, CI, test automation, and ephemeral environment platform for a large TypeScript, React, and Go monorepo. The role requires 6+ years of large-scale build-system experience, strong Bazel or comparable tooling expertise, and deep knowledge of hermetic, reproducible development workflows.
Senior platform engineer responsible for reliable, secure, and scalable infrastructure, developer tooling, observability, and AI enablement. The role requires 6+ years in platform engineering, SRE, or DevOps, with strong AWS and incident leadership experience.
Own reliability, scalability, security, observability, and incident response for mission-critical applications across Kubernetes, AWS, and on-premise DoD environments. Requires an active Top Secret clearance and at least five years of infrastructure-focused SRE, DevOps, or platform engineering experience.
Build and operate production software, APIs, and automation for large-scale bare-metal and GPU infrastructure. The role requires 8+ years of software or infrastructure engineering experience, strong Python and Linux skills, and expertise in provisioning, lifecycle management, and reliability.