Senior Platform Engineer: Storage
Designs and evolves production Ceph storage clusters, builds APIs and orchestration services for block/object storage using Go and gRPC. Requires experience with distributed systems, filesystems like ZFS/BTRFS, and building scalable infrastructure.
About the job
Responsibilities
- Design and evolve multiple production Ceph clusters, from hardware design, to driving network requirements to configuring, tuning and operating clusters and their clients
- Create efficient, generalizable APIs using systems/kernel features to provide safe, as-fast-as-possible live-migrations of stateful workload between hosts
- Design and build API and Orchestration services to tie storage primitives to higher level primitives using Go, gRPC, ScyllaDB and Temporal
- Write Engineering Requirement Documents to take something from idea, to defined tasks, to implementation, to monitoring it’s success
- Design build a suite of storage primitives that can be used by customer applications, internal services and enable higher level platform features such as streaming image pulls or movable build caches
Requirements
- Experience architecting and implementing distributed systems. You enjoy building fault tolerant, resilient, and scalable services
- Production experience with distributed block device systems (Ceph) or a solid understanding of network storage cluster design from first principles
- Understanding and experience with current gen filesystems (Ext4, ZFS, BTRFS). Bonus points for next gen (EROFS, bcachefs)
- A solid intuition about how long your solutions will last. All systems age. In startups, we can hope for 2-3 orders of magnitude, or 12-18mo
- The tact to implement your solution, creator monitors for it’s error boundaries, and document any requirements for when you’re not around
- A great sense of direction and prioritization when it comes to dealing with the ambiguity of an early stage startup
- A sense of grit to dive into a problem, implement a solution, scale that solution, and replace it when needed
- A great set of communication skills for getting your point across, solution implemented, and beyond
Skills
Ceph, Go, gRPC, Scylladb, Temporal, Zfs, Btrfs, Ext4, Linux Kernel, Distributed Systems
Similar jobs
DevOps / SRE jobsDesigns and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.
Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.
Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.
Build and mature Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes optimization, environment bootstrapping, and cost optimization. The role requires 5+ years of software engineering experience, cloud-native expertise, and strong technical leadership.
Senior Software Engineer building and improving Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes, cloud optimization, and developer productivity workflows. Requires 5+ years of software engineering experience and expertise in cloud-native or platform engineering.