Skip to content
RunpodRunpod

Director of Infrastructure Engineering

Lead and scale Runpod's core cloud and bare-metal infrastructure, including SRE, global/HPC networking, and distributed storage for massive GPU/AI workloads. Requires 8+ years operating large-scale distributed systems plus 7+ years leading infrastructure/SRE teams (managing managers).

About the job

Responsibilities

  • Own Core Infrastructure & SRE: Lead multiple engineering teams responsible for Site Reliability Engineering, networking, and storage. Establish rigorous SRE practices, driving SLA/SLO definitions, incident response, observability, and automated remediation.
  • Architect HPC & Global Networking: Oversee the design, scaling, and operation of Runpod’s global network backbone, as well as ultra-low-latency HPC cluster networks. Drive the implementation and optimization of InfiniBand and RDMA over Converged Ethernet (RoCE) to support massive, multi-node GPU training workloads.
  • Drive Storage Engine Innovation: Direct the architecture and performance tuning of highly scalable, distributed storage systems. Ensure our storage engines can deliver the massive IOPS and throughput required to keep high-end GPUs fed with data during deep learning tasks.
  • Build a High-Output Org: Hire, mentor, and grow highly technical engineering managers and senior ICs (network architects, systems engineers, SREs). Create a culture of ownership, operational excellence, and craft in a remote-first environment.
  • Translate Scale into Strategy: Partner with Program Management and Product to forecast capacity requirements, shape technical roadmaps, and convert massive scale challenges into clear technical scopes, milestones, and measurable outcomes.
  • Continuously Improve Systems & Flow: Drive measurable improvements in infrastructure reliability and delivery metrics, such as deployment frequency, MTTR (Mean Time To Recovery), infrastructure as code (IaC) coverage, and system uptime.
  • Architectural Stewardship: Provide architectural oversight for bare-metal provisioning, virtualization layers, network fabrics, and storage clusters, ensuring seamless scalability without becoming a bottleneck for your teams.
  • Cross-Functional Partnership: Coordinate cleanly with product delivery and platform teams to ensure the infrastructure primitives they rely on are robust, well-documented, and highly available.

Requirements

  • 7+ years leading software, infrastructure, SRE, or networking teams, including managing managers and multiple squads, with a proven record of scaling high-availability cloud environments.
  • 8+ years building and operating large-scale distributed systems, bare-metal infrastructure, or public/private cloud platforms.
  • Proven hands-on background or strong architectural understanding of ultra-low latency networking. Deep familiarity with InfiniBand and/or RoCE, spine-leaf architectures, and global WAN routing protocols (BGP).
  • Experience building, operating, or tuning high-performance distributed storage systems and parallel file systems (e.g., Ceph, Lustre, Weka, NVMe-oF) capable of handling heavy AI/ML I/O loads.
  • Strong foundation in reliability engineering, infrastructure-as-code (Terraform, Ansible), container orchestration (Kubernetes), and modern observability stacks.
  • Experience building culture, accountability, and momentum across distributed technical teams.
  • Clear written and verbal communication, strong stakeholder management, and calm, decisive leadership during high-stakes operational incidents.

Preferred Qualifications

  • Direct experience architecting and operating infrastructure specifically optimized for massive GPU clusters and AI/ML workloads.
  • Deep understanding of hardware architectures, GPU interconnects (NVLink), and datacenter topology.
  • Track record of scaling infrastructure teams in hyper-growth startup environments.
  • Open-source contributions or active recognition within the infrastructure, networking, or Kubernetes communities.

Compensation & Benefits

  • Competitive base pay ranges from $225,000 - $325,000 (may be inclusive of several career levels; narrowed during interview process based on experience, qualifications, and location).
  • Meaningful equity in a fast-growing company (stock options for everyone).
  • Generous medical, dental & vision plans.
  • Flexible PTO.
  • $1,200 Home Office & Equipment Stipend.
  • Remote-first with Slack as main communication tool.

Skills

SRE, Hpc, InfiniBand, Roce, Kubernetes, Terraform, Ansible, BGP, Ceph, Lustre, Nvme-Of, Observability, Distributed Systems, Gpu Clusters

Bestow

Bestow

United States

Engineering Director
$225k+/yrRemote8+ YOEEngineering Management

Leads multiple engineering teams responsible for insurance product configuration, enrollment workflows, and core domain services. The role requires deep distributed-systems expertise, backend development experience, strong operational leadership, and a track record of managing managers in a regulated or enterprise environment.

Motive

Motive

San Francisco, CA
Director, Developer Platform & Experience
$229k+/yrHybrid12+ YOEEngineering Management

Leads a multi-team organization responsible for AI-assisted development, cloud agentic infrastructure, developer experience, CI/CD, testing, and engineering velocity. Requires extensive software engineering and engineering leadership experience, deep infrastructure expertise, and hands-on knowledge of AI developer tooling and LLM evaluation.

Rippling

Rippling

San Francisco, CA

Director of Engineering, HRIS Core Flows
$216k+/yrOn-site8+ YOEEngineering Management

Leads a 15-person engineering organization responsible for enterprise employee lifecycle workflows and an AI-assisted HR workflow charter. The role requires multi-team leadership, strong technical and product judgment, and experience with high-stakes, global, workflow-heavy platforms.

Okta

Okta

San Francisco, CA

Senior Director, GRC
$236k+/yrOn-site10+ YOEEngineering Management

Leads Okta’s security GRC organization, overseeing enterprise cyber risk, AI governance, global compliance, audits, vendor risk, and engineering-driven remediation. The role requires 10+ years of progressive Security GRC leadership, cloud technology experience, AI governance expertise, and a bachelor’s degree or equivalent experience.

LeafLink

LeafLink

United States

Head of Engineering
$240k+/yrRemote12+ YOEEngineering Management

Leads the entire engineering organization, owning technical vision, architecture, execution standards, security, budget, and organizational scaling. The role requires extensive software engineering and engineering management experience, cloud-native architecture expertise, and a record of building high-performing teams.