Skip to content
The Voleon GroupThe Voleon GroupBerkeley, CA

Storage and Datacenter Team Lead

Leads a storage engineering team while architecting and operating highly available Linux-based storage, datacenter, and data-protection infrastructure. The role requires deep Ceph experience, PB-scale archiving and backup expertise, hands-on troubleshooting, and team leadership.

215k – 245k/yr
Remote5+ YOEDevOps / SRE

About the role

Responsibilities

Leadership & Team Management

  • Lead a small team of storage, database, and systems administrators, including mentorship, performance management, and career development.
  • Coordinate datacenter operations across production and research facilities, including site-work scheduling and vendor or contractor visits.
  • Align team priorities with organizational goals and ensure timely project delivery.
  • Participate in hiring to grow and evolve the storage engineering team.
  • Coordinate on-call schedules and maintain effective incident-response processes.

Technical & Operational Oversight

  • Architect, implement, and maintain highly available, performant storage systems.
  • Define and drive automation strategies for storage deployment and monitoring.
  • Oversee storage lifecycle management, including capacity planning, performance tuning, archiving, backups, and data protection for large-scale datasets.
  • Provide architectural guidance and hands-on support for Ceph at PB scale.
  • Oversee physical datacenter infrastructure, including rack layout, power distribution, cooling systems, and space, power, and cooling capacity forecasting.
  • Manage equipment installation and decommissioning.
  • Collaborate with networking, virtualization, research, and application teams to support compute and storage needs.
  • Improve CI/CD and configuration-management processes using tools such as Ansible and Git.
  • Support database operations through database tuning, storage optimization, and collaboration with developers.
  • Develop runbooks for remote-hands work and coordinate onsite operations with contractors and facilities personnel.

IC-Level Engineering & Troubleshooting

  • Serve as an escalation point for distributed filesystems, databases, and high-performance storage infrastructure.
  • Administer Linux servers, network-attached storage, virtualization platforms, and cluster frameworks.
  • Install, cable, and troubleshoot physical server, storage, and network hardware in rack environments.
  • Diagnose and resolve hardware-level issues affecting production systems.
  • Support and enhance observability using Prometheus, Grafana, and related tools.

Requirements

  • 5+ years of Linux systems administration experience, with significant recent focus on storage systems.
  • 2+ years of team leadership, technical project management, or mentoring experience.
  • Knowledge of distributed storage systems such as Ceph and storage technologies including RAID, SAN, and NAS.
  • Experience streamlining data lifecycle processes, including archiving, backup, and retention of PB-scale data.
  • Hands-on experience with colocated datacenter infrastructure.
  • Ability to travel to remote datacenter sites as needed.
  • Strong scripting or development experience in Bash and/or Python.
  • Experience with Ansible and infrastructure automation.
  • Familiarity with monitoring and alerting systems such as Nagios, CheckMK, Prometheus, and Grafana.
  • Understanding of virtualization technologies including KVM and ESXi, and containerization technologies including Docker and Podman.
  • Knowledge of LDAP, IPA, AD, and centralized identity management.

Preferred Qualifications

  • Experience with Kubernetes container orchestration.
  • PostgreSQL DBA experience.
  • Experience in a high-throughput research or trading environment.
  • Exposure to RHEL, CentOS, or Rocky Linux in enterprise settings.
  • Experience with DCIM tools for tracking assets, power, and space.
  • Familiarity with CI/CD pipelines and DevOps principles.
  • Experience managing colocation vendor relationships and SLAs.

Skills

LinuxcephraidsannasBashPythonAnsiblePrometheusGrafanakvmesxiDockerpodmanKubernetes

Similar roles

DevOps / SRE jobs
OfferUp

Senior Cloud Engineer

OfferUpBellevue, WA +1

Senior Cloud Engineer owning AWS/GCP infrastructure, Kubernetes/GitOps platforms, and CI/CD systems. Designs and operates scalable, secure cloud infrastructure while mentoring engineers and enabling AI/ML tooling.

215k – 240k/yrHybrid5+ YOEDevOps / SRE
Nooks

Senior Software Engineer, Core Infrastructure

NooksSan Francisco, CA

Senior engineer on the Core Infrastructure team responsible for scaling data layers, observability, and developer tooling to support rapid multi-product growth at Nooks. Requires 5+ years experience scaling systems 10x+, strong distributed systems or infra background, and willingness to be in-office in San Francisco 3+ days/week.

215k – 300k/yrHybrid5+ YOEDevOps / SRE
The Voleon Group

Senior Linux Infrastructure Engineer

The Voleon GroupUnited States

Builds, automates, and maintains scalable, secure Linux infrastructure using modern tooling. Requires 7+ years Linux sysadmin experience, deep OS/network knowledge, Python/Ruby scripting, config management (Ansible/Puppet/Chef), and monitoring (Prometheus/Grafana).

215k – 250k/yrRemote7+ YOEDevOps / SRE
Carta

Senior Software Engineer II, Developer Experience

CartaSan Francisco, CA +2

Builds AI-native developer tooling including MCP servers, agents, and CI/CD pipelines to enhance Carta's software delivery lifecycle. Requires 8+ years experience shipping production LLM systems and platform engineering with Python/Java, cloud-native tech.

213k – 250k/yrOn-site8+ YOEDevOps / SRE
Tines

Senior Site Reliability Engineer - Government Cloud

TinesUnited States

Build and operate AWS GovCloud infrastructure for federal customers, owning IaC, container pipelines, compliance documentation, and operational tooling. Requires 5+ years AWS experience and FedRAMP familiarity.

210k – 220k/yrRemote5+ YOEDevOps / SRE