Infrastructure Engineer (Storage)
Operate and scale distributed storage systems like VAST and Ceph for AI/ML workloads, build Python automation tools, manage Linux bare-metal systems, and collaborate on infrastructure optimizations. Requires 5+ years in infrastructure engineering with storage expertise.
About the job
What You'll Do
Storage Systems & Infrastructure
- Operate and scale distributed storage systems, including VAST and S3-compatible object storage (e.g., Ceph)
- Improve performance, reliability, and efficiency of storage systems supporting large-scale AI/ML workloads
- Troubleshoot complex storage and data path issues across hardware and software layers
- Optimize storage performance to support high-throughput, low-latency AI training and inference workloads
Automation & Tooling
- Build and maintain automation for provisioning, managing, and monitoring storage infrastructure
- Develop Python-based tools and workflows to reduce manual operational overhead
- Improve lifecycle management of storage clusters, from deployment through maintenance and scaling
Systems & Operations
- Manage and operate Linux-based systems in production, including bare-metal environments
- Partner with infrastructure and data center teams on hardware bring-up, upgrades, and issue resolution
- Support capacity planning, utilization tracking, and forecasting for storage systems
- Leverage monitoring and telemetry to diagnose issues and improve system performance and reliability
Cross-Functional Collaboration
- Work closely with Infrastructure Engineering, Network Engineering, and Platform teams to integrate storage into the broader platform
- Contribute to design discussions around new infrastructure deployments and scaling strategies
- Help define best practices for operating storage systems in high-performance computing environments
What You'll Need
Required Qualifications
- 5+ years of experience in infrastructure engineering, systems engineering, or related roles
- Hands-on experience operating distributed storage systems (e.g., VAST, Ceph, or similar)
- Strong Linux systems experience in production environments
- Proficiency in Python or similar scripting/programming languages for automation
- Experience working with bare-metal infrastructure and hardware-oriented systems
- Ability to debug complex issues across system boundaries (storage, OS, hardware, networking)
- Experience with storage networking protocols (e.g., NFS or similar)
- Experience with capacity planning, monitoring, and performance tuning
Ideal Experience
- Experience with VAST storage systems in production environments
- Experience operating S3-compatible object storage at scale
- Data center operations experience, including working with physical hardware
- Familiarity with AI/ML or HPC workloads and their storage requirements
- Background in high-performance or low-latency distributed systems
- Familiarity with high-performance data transfer technologies (e.g., RDMA, GPU Direct Storage)
- Experience supporting GPU-based workloads or large-scale compute clusters
Compensation
Anticipated annual base salary range: $180,000—$200,000 USD
Skills
Vast, Ceph, Linux, Python, Bare-Metal, Nfs, S3, Rdma, Gpu Direct Storage, Monitoring
Similar jobs
DevOps / SRE jobsBuild developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.
Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.
Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.