# Storage and Datacenter Team Lead

**Company:** [The Voleon Group](https://hotfix.jobs/companies/voleon)
**Location:** Remote
**Role:** DevOps / SRE
**Salary:** $215k – $245k/yr
**Experience:** 5+ years
**Skills:** Linux, ceph, raid, san, nas, Bash, Python, Ansible, Prometheus, Grafana, kvm, esxi, Docker, podman, Kubernetes
**Posted:** 2026-08-04

> Leads a storage engineering team while architecting and operating highly available Linux-based storage, datacenter, and data-protection infrastructure. The role requires deep Ceph experience, PB-scale archiving and backup expertise, hands-on troubleshooting, and team leadership.

## Job Description

## Responsibilities

### Leadership & Team Management
- Lead a small team of storage, database, and systems administrators, including mentorship, performance management, and career development.
- Coordinate datacenter operations across production and research facilities, including site-work scheduling and vendor or contractor visits.
- Align team priorities with organizational goals and ensure timely project delivery.
- Participate in hiring to grow and evolve the storage engineering team.
- Coordinate on-call schedules and maintain effective incident-response processes.

### Technical & Operational Oversight
- Architect, implement, and maintain highly available, performant storage systems.
- Define and drive automation strategies for storage deployment and monitoring.
- Oversee storage lifecycle management, including capacity planning, performance tuning, archiving, backups, and data protection for large-scale datasets.
- Provide architectural guidance and hands-on support for Ceph at PB scale.
- Oversee physical datacenter infrastructure, including rack layout, power distribution, cooling systems, and space, power, and cooling capacity forecasting.
- Manage equipment installation and decommissioning.
- Collaborate with networking, virtualization, research, and application teams to support compute and storage needs.
- Improve CI/CD and configuration-management processes using tools such as Ansible and Git.
- Support database operations through database tuning, storage optimization, and collaboration with developers.
- Develop runbooks for remote-hands work and coordinate onsite operations with contractors and facilities personnel.

### IC-Level Engineering & Troubleshooting
- Serve as an escalation point for distributed filesystems, databases, and high-performance storage infrastructure.
- Administer Linux servers, network-attached storage, virtualization platforms, and cluster frameworks.
- Install, cable, and troubleshoot physical server, storage, and network hardware in rack environments.
- Diagnose and resolve hardware-level issues affecting production systems.
- Support and enhance observability using Prometheus, Grafana, and related tools.

## Requirements
- 5+ years of Linux systems administration experience, with significant recent focus on storage systems.
- 2+ years of team leadership, technical project management, or mentoring experience.
- Knowledge of distributed storage systems such as Ceph and storage technologies including RAID, SAN, and NAS.
- Experience streamlining data lifecycle processes, including archiving, backup, and retention of PB-scale data.
- Hands-on experience with colocated datacenter infrastructure.
- Ability to travel to remote datacenter sites as needed.
- Strong scripting or development experience in Bash and/or Python.
- Experience with Ansible and infrastructure automation.
- Familiarity with monitoring and alerting systems such as Nagios, CheckMK, Prometheus, and Grafana.
- Understanding of virtualization technologies including KVM and ESXi, and containerization technologies including Docker and Podman.
- Knowledge of LDAP, IPA, AD, and centralized identity management.

## Preferred Qualifications
- Experience with Kubernetes container orchestration.
- PostgreSQL DBA experience.
- Experience in a high-throughput research or trading environment.
- Exposure to RHEL, CentOS, or Rocky Linux in enterprise settings.
- Experience with DCIM tools for tracking assets, power, and space.
- Familiarity with CI/CD pipelines and DevOps principles.
- Experience managing colocation vendor relationships and SLAs.

## Similar roles

- [Senior Cloud Engineer](https://hotfix.jobs/jobs/40ff753a-7206-4118-9197-ad9136840d10) - OfferUp - Bellevue, WA - $215k – $240k/yr
- [Senior Software Engineer, Core Infrastructure](https://hotfix.jobs/jobs/3bf81467-bf7d-43a0-9e03-975181234049) - Nooks - San Francisco, CA - $215k – $300k/yr
- [Senior Linux Infrastructure Engineer](https://hotfix.jobs/jobs/f025b639-9cac-4c4b-8a3e-b57a0c093e56) - The Voleon Group - Remote - $215k – $250k/yr
- [Senior Software Engineer II, Developer Experience](https://hotfix.jobs/jobs/affd1bbf-2289-453a-a464-026ad0946432) - Carta - San Francisco, CA - $213k – $250k/yr
- [Senior Site Reliability Engineer - Government Cloud](https://hotfix.jobs/jobs/2baea467-f6d7-4c95-8eda-8bf4a56ab384) - Tines - Remote - $210k – $220k/yr

**Apply:** https://hotfix.jobs/jobs/367f1d16-228b-4832-8a2a-6ee32afbf54f
**Canonical:** https://hotfix.jobs/jobs/367f1d16-228b-4832-8a2a-6ee32afbf54f