Senior Software Engineer, Storage
Designs, builds, and optimizes distributed cloud storage systems for AI/HPC workloads, focusing on high-performance filesystems, block/object storage, and Linux subsystems. Requires deep expertise in scalable storage infrastructure and languages like Go, C++, Rust.
About the job
What You’ll Be Working On
Building Our Multi-Petabyte Cloud Storage Platform
- Building core components of our foundational storage products, purpose built for high performance AI and ML workloads
- Contributing to distributed file, block and object storage products, with a focus on filesystem based solutions
System Design & Architecture
- Design and implement high-performance, scalable, and resilient storage architectures that are highly extensible
- Proposing and prototyping novel strategies to scale performance and system throughput for our most demanding customer workloads
- Building observability, metrics and tooling for our services and fleet
High Velocity Problem Solving
- Troubleshooting and resolving unique and complex distributed systems problems only seen at the scale we operate at
- Provide ongoing support for production systems, and customer workloads including troubleshooting, performance tuning, and incident response
Cross-functional Collaboration
- Foster strong collaboration with other engineering teams (e.g., Software Infrastructure, Product) and cross-functional departments
- Single threaded ownership and representation of the storage team in business critical initiatives across the company
You’ll Bring to the Team
- Hands-on proficiency in modern software development best practices, and practical experience in languages like Go, Java, C/C++, or Rust
- Extensive experience developing multi-tenant, cloud scale distributed storage infrastructure software and systems
- Experience contributing to at least one or more of the following storage products: File (e.g., NFS, SMB, Lustre), Object, or Block Storage (e.g., NVMe, iSCSI)
- A strong background in high performance filesystem based products, VFS and linux filesystems (e.g., ext4, XFS, ZFS)
- Proficiency working with Linux and its storage subsystems
- Knowledge of monitoring tools (Prometheus, Grafana), log analysis, distributed tracing and debugging
Bonus Points
- Experience with AI/HPC storage solutions, such as Parallel Filesystems or petabyte+ scale Object Storage
- Familiarity with networking technologies like RDMA and Infiniband
- Familiarity with modern storage technologies (e.g GPU Direct Storage, F2FS, SPDK etc)
- Prior experience with Nvidia SuperNIC DPUs for storage optimization
- Prior experience in Storage Virtualization & Orchestration, volume placement strategies and distributed metadata management
- Research publications or open-source contributions to storage-related projects
Compensation
Compensation will be paid in the range of $166,000 - $201,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant’s education, experience, knowledge, skills, and abilities, as well as internal equity and alignment with market data.
Skills
Go, Java, C++, Rust, Linux, Nfs, Smb, Lustre, Nvme, Iscsi, Ext4, Xfs, Zfs, Vfs, Prometheus
Similar jobs
Backend Engineering jobsSenior Software Engineer who will architect and scale distributed backend data systems supporting patient services. The role requires 4+ years of experience, expertise in microservices and event-driven systems, and strong technical leadership and mentoring skills.
Build and operate scalable backend platforms, data pipelines, and services that create real-time digital twins of store inventory. The role requires 5+ years of software engineering experience, backend programming expertise, distributed systems knowledge, and cloud infrastructure experience.
Build and operate production fintech and payment services, APIs, and developer tooling using Go and microservices. The role requires experience delivering complex production systems, working in payments, and maintaining strong standards for testing, observability, and AI-assisted development.
Design, build, and operate Cloudflare’s globally distributed cache and reverse-proxy data plane, improving performance, correctness, and resilience across the edge. Requires at least 4 years of production systems experience and proficiency in a systems or backend language.
Senior engineer building secure, scalable backend services and cross-stack product features for Rippling’s emerging HR, IT, Finance, and Spend products. The role requires 5+ years of production experience with Python, Django, or related technologies, plus frontend familiarity with React or TypeScript.