Staff Software Engineer, Storage
Leads architectural vision and strategy for AI-scale cloud storage systems, bridging NVMe hardware to distributed object stores. Requires 12+ years in distributed systems, system programming in C/C++/Rust/Go, and deep storage protocol expertise.
About the job
What You’ll Be Working On
Architectural Vision & Strategy: Define and drive the long-term technical strategy for Crusoe’s storage engine. Identify industry trends (e.g., CXL, NVMe-oF) and integrate them into a cohesive roadmap.
System Programming Expertise: Leverage proven experience in system programming with languages such as C, C++, Go, and/or Rust to build the foundations of our V2 storage re-architecture.
Storage Protocols: Architect and implement solutions utilizing industry-standard storage protocols such as NFS, SMB, iSCSI, and NVMe/TCP.
Open Source Stewardship: Drive and maintain a track record of contributions to the open-source community (e.g., Ceph, GlusterFS, Lustre, Spectrum Scale, OpenEBS).
Technical Authority: Serve as the final arbiter for critical architecture decisions across the Foundations organization. Lead complex design reviews that intersect storage, networking, and virtualization.
Deep Performance Engineering: Lead "tiger teams" to solve the most ambiguous and difficult bottlenecks in the stack—from kernel-level IO context switching to global tail-latency in distributed clusters.
Strategic Collaboration: Work closely with Executive Leadership and Product to align technical capabilities with business milestones.
What You’ll Bring to the Team
Cloud Storage Expertise: 12+ years of experience building and operating large-scale, complex distributed cloud computing infrastructure products.
Troubleshooting & Tuning: Strong troubleshooting and performance tuning skills; ability to profile and optimize the entire IO path.
High-Drive Mindset: Self-motivation to thrive in a fast-paced environment with a high degree of ownership and minimal supervision.
Masters of Consistency & Durability: Deep theoretical and practical knowledge of distributed state and data protection at petabyte scale.
Software Engineering Fundamentals: Mastery of professional software engineering practices for the full SDLC, including coding standards, build processes, and testing.
Communication & Collaboration: Ability to champion and lead initiatives across the engineering organization, such as tech talks and technical reading groups.
Bonus Points
- Public Cloud & AI/ML: Expertise in one or more Public Cloud offerings (AWS, GCP, Azure, OCI) and familiarity with AI/ML frameworks (PyTorch, Tensorflow, JAX) and MLOps.
- High-Throughput I/O: Experience with cutting-edge I/O architectures like DAOS or SPDK.
- Networking Foundations: Background in RDMA and high-performance networking, including SmartNICs and RoCEv2.
- Distributed Systems Mastery: Experience with highly available and scalable systems such as Cassandra, MongoDB, Redis, or Kafka.
- Theoretical Depth: Strong knowledge of distributed systems fundamentals including CAP Theorem, Paxos/RAFT, consistent hashing, and sharding strategies.
Education: Advanced degree (Master's or PhD) in Computer Science, Engineering, or a related field.
Benefits
- Competitive compensation
- Restricted Stock Units
- Paid time off & paid holidays
- Comprehensive health, dental & vision insurance
- Employer contributions to HSA account
- Paid parental leave
- Paid life insurance, short-term and long-term disability
- Professional development & tuition reimbursement
- Mental health & wellness support
- Commuter benefits (parking & transit)
- Cell phone stipend
- 401(k) Retirement plan with company match up to 4% of salary
- Volunteer time off
Compensation Range: Up to $240,000 - $310,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.
Skills
C++, Rust, Go, Ceph, Nvme, Nfs, Iscsi, Nvme/Tcp, Cxl, Nvme-Of, Spdk, Daos, Rdma, Rocev2, AWS
Similar jobs
Backend Engineering jobsStaff Engineer responsible for the long-term technical health and strategic direction of Carta’s Middle Office domain, including distributed systems, architecture, reliability, and AI-enabled engineering practices. Requires 10+ years of software engineering experience and demonstrated technical leadership across teams.
Staff Software Engineer on Snowflake’s Snowtrail team, designing and operating query-processing and distributed infrastructure that improves engineering velocity and production quality. Requires expertise in databases, query optimization, compiler implementation, and large-scale systems.
Leads the design and implementation of scalable distributed data pipelines and related infrastructure, while improving performance, reliability, and security. Requires 5+ years of software experience, systems programming expertise, cloud and containerization experience, and senior staff-level technical leadership.
Leads performance engineering for high-volume data systems by diagnosing bottlenecks, designing improvements, and building resilient performance platforms. Requires 10+ years of software experience, strong Java and database expertise, and experience with distributed backend systems.
Senior technical IC focused on backend architecture and AI technologies for Airbnb's communication and connectivity platform. Partners with senior leaders and contributes code while providing technical leadership across teams.