Director of Infrastructure Engineering
Lead and scale Runpod's core cloud and bare-metal infrastructure, including SRE, global/HPC networking, and distributed storage for massive GPU/AI workloads. Requires 8+ years operating large-scale distributed systems plus 7+ years leading infrastructure/SRE teams (managing managers).
About the job
Responsibilities
- Own Core Infrastructure & SRE: Lead multiple engineering teams responsible for Site Reliability Engineering, networking, and storage. Establish rigorous SRE practices, driving SLA/SLO definitions, incident response, observability, and automated remediation.
- Architect HPC & Global Networking: Oversee the design, scaling, and operation of Runpod’s global network backbone, as well as ultra-low-latency HPC cluster networks. Drive the implementation and optimization of InfiniBand and RDMA over Converged Ethernet (RoCE) to support massive, multi-node GPU training workloads.
- Drive Storage Engine Innovation: Direct the architecture and performance tuning of highly scalable, distributed storage systems. Ensure our storage engines can deliver the massive IOPS and throughput required to keep high-end GPUs fed with data during deep learning tasks.
- Build a High-Output Org: Hire, mentor, and grow highly technical engineering managers and senior ICs (network architects, systems engineers, SREs). Create a culture of ownership, operational excellence, and craft in a remote-first environment.
- Translate Scale into Strategy: Partner with Program Management and Product to forecast capacity requirements, shape technical roadmaps, and convert massive scale challenges into clear technical scopes, milestones, and measurable outcomes.
- Continuously Improve Systems & Flow: Drive measurable improvements in infrastructure reliability and delivery metrics, such as deployment frequency, MTTR (Mean Time To Recovery), infrastructure as code (IaC) coverage, and system uptime.
- Architectural Stewardship: Provide architectural oversight for bare-metal provisioning, virtualization layers, network fabrics, and storage clusters, ensuring seamless scalability without becoming a bottleneck for your teams.
- Cross-Functional Partnership: Coordinate cleanly with product delivery and platform teams to ensure the infrastructure primitives they rely on are robust, well-documented, and highly available.
Requirements
- 7+ years leading software, infrastructure, SRE, or networking teams, including managing managers and multiple squads, with a proven record of scaling high-availability cloud environments.
- 8+ years building and operating large-scale distributed systems, bare-metal infrastructure, or public/private cloud platforms.
- Proven hands-on background or strong architectural understanding of ultra-low latency networking. Deep familiarity with InfiniBand and/or RoCE, spine-leaf architectures, and global WAN routing protocols (BGP).
- Experience building, operating, or tuning high-performance distributed storage systems and parallel file systems (e.g., Ceph, Lustre, Weka, NVMe-oF) capable of handling heavy AI/ML I/O loads.
- Strong foundation in reliability engineering, infrastructure-as-code (Terraform, Ansible), container orchestration (Kubernetes), and modern observability stacks.
- Experience building culture, accountability, and momentum across distributed technical teams.
- Clear written and verbal communication, strong stakeholder management, and calm, decisive leadership during high-stakes operational incidents.
Preferred Qualifications
- Direct experience architecting and operating infrastructure specifically optimized for massive GPU clusters and AI/ML workloads.
- Deep understanding of hardware architectures, GPU interconnects (NVLink), and datacenter topology.
- Track record of scaling infrastructure teams in hyper-growth startup environments.
- Open-source contributions or active recognition within the infrastructure, networking, or Kubernetes communities.
Compensation & Benefits
- Competitive base pay ranges from $225,000 - $325,000 (may be inclusive of several career levels; narrowed during interview process based on experience, qualifications, and location).
- Meaningful equity in a fast-growing company (stock options for everyone).
- Generous medical, dental & vision plans.
- Flexible PTO.
- $1,200 Home Office & Equipment Stipend.
- Remote-first with Slack as main communication tool.
Skills
SRE, Hpc, InfiniBand, Roce, Kubernetes, Terraform, Ansible, BGP, Ceph, Lustre, Nvme-Of, Observability, Distributed Systems, Gpu Clusters
Similar jobs
Engineering Management jobsLeads multiple engineering teams responsible for insurance product configuration, enrollment workflows, and core domain services. The role requires deep distributed-systems expertise, backend development experience, strong operational leadership, and a track record of managing managers in a regulated or enterprise environment.
Leads a multi-team organization responsible for AI-assisted development, cloud agentic infrastructure, developer experience, CI/CD, testing, and engineering velocity. Requires extensive software engineering and engineering leadership experience, deep infrastructure expertise, and hands-on knowledge of AI developer tooling and LLM evaluation.
Leads a 15-person engineering organization responsible for enterprise employee lifecycle workflows and an AI-assisted HR workflow charter. The role requires multi-team leadership, strong technical and product judgment, and experience with high-stakes, global, workflow-heavy platforms.
Leads Okta’s security GRC organization, overseeing enterprise cyber risk, AI governance, global compliance, audits, vendor risk, and engineering-driven remediation. The role requires 10+ years of progressive Security GRC leadership, cloud technology experience, AI governance expertise, and a bachelor’s degree or equivalent experience.
Leads the entire engineering organization, owning technical vision, architecture, execution standards, security, budget, and organizational scaling. The role requires extensive software engineering and engineering management experience, cloud-native architecture expertise, and a record of building high-performing teams.