Skip to content
RunpodRunpodUnited States

Datacenter Infrastructure Specialist

Own the technical lifecycle and operational health of Runpod’s global high-density GPU fleet. Bridge hardware partners and engineering teams through hardware validation, network troubleshooting, AI-driven automation, incident response, and performance tuning for AI/ML workloads.

Salary not listed
Remote3+ YOEDevOps / SRE

About the role

Responsibilities

  • Assist in validating new hardware, ensuring partner deployments meet Runpod’s specifications for distributed AI/ML workloads.
  • Monitor fleet health to identify performance degradation. Help audit downtime and provide technical data to protect customer SLAs.
  • Work with LLMs and AI agents to automate network triage and generate dynamic runbooks for the fleet.
  • Coordinate technical incident communications with clear updates, translating outages into actionable resolutions.
  • Support the growth of infrastructure partners as the technical authority bridging hardware partners and internal engineering teams.
  • Own the technical lifecycle and operational health of the high-density GPU fleet, including HPC systems engineering, advanced network troubleshooting, and process automation.

Requirements

  • 3–5 years of experience in infrastructure operations, systems reliability, or datacenter engineering.
  • Strong proficiency in datacenter networking and performance troubleshooting.
  • Hands-on experience with the NVIDIA Software Stack (driver installation, performance utilities) and understanding of multi-node performance tuning.
  • Solid Linux system administration skills and experience with containerization (Docker).
  • Comfortable with system-level troubleshooting and performance tuning at the kernel and hardware interface layers.
  • Clear written and verbal communication skills to explain hardware or networking issues to technical partners and internal leadership.
  • Detail-oriented and proactive in identifying potential failures before they impact customers.
  • Willingness to participate in an on-call rotation as the global fleet scales.

Nice-to-Haves

  • Experience working in a fast-paced startup environment contributing to building operational workflows.
  • Experience managing or optimizing bare-metal High-Performance Computing environments at massive scale.
  • Experience with observability tools such as Grafana, Prometheus, or Datadog.
  • Proficiency in Python, Go (Golang), or Bash to automate repetitive infrastructure tasks and interface with internal APIs.
  • Exposure to RDMA, InfiniBand, or RoCE.

Skills

nvidia software stackLinuxDockerrdmaInfiniBandroceGrafanaPrometheusDatadogPythonGoBashhpc

Similar roles

DevOps / SRE jobs
Writer

Infrastructure engineer

WriterNew York, NY +2

Infrastructure Engineer building and operating scalable, reliable production systems for an enterprise AI platform. Owns end-to-end reliability, automates with Python/Go, integrates AI agents into workflows, leads incident response, and collaborates cross-functionally on high-availability infrastructure using Kubernetes, Terraform, and multi-cloud tooling. Requires 5+ years experience and daily use of AI tooling.

140k – 274k/yrHybrid5+ YOEDevOps / SRE
ConductorOne

Site Reliability Engineer

ConductorOnePortland, ME

Own reliability and scalability of a horizontal identity platform, including core cloud and FedRAMP environments. Build observability, automate operations, lead incident response, and ensure new features are reliable from the start. Requires production SRE experience at scale with Kubernetes, IaC, and strong programming skills.

180k – 250k/yrHybrid5+ YOEDevOps / SRE
Closinglock

DevOps Engineer

ClosinglockAustin, TX

DevOps Engineer building and maintaining scalable AWS cloud infrastructure, CI/CD pipelines, IaC, observability, and Kubernetes workloads for a security-critical fintech platform. Requires 3-5 years DevOps/SRE experience, strong AWS and infrastructure skills, on-call participation.

Salary not listedOn-site3+ YOEDevOps / SRE
Tulip

AI Enablement Engineer

TulipSomerville, MA

Build and maintain an internal agentic AI platform to accelerate developer workflows at Tulip. Identify high-impact AI opportunities in code generation, testing, debugging and tooling; own evals, standards, onboarding and measurement of AI adoption. Requires 5+ years software engineering experience with strong hands-on LLM/agentic AI and full-stack TypeScript skills.

130k – 180k/yrHybrid5+ YOEDevOps / SRE
OpenAI

Network Operations Engineer, Network Activation

OpenAISan Francisco, CA

Own hands-on Layer 1 activation, troubleshooting, and automation for WAN, fiber, carrier, and cloud interconnect circuits in OpenAI's GPU fleet data centers. Requires 3+ years network operations experience, optical fault isolation skills, provider coordination, and practical automation with scripts/APIs/LLM tools.

157k – 221k/yrOn-site3+ YOEDevOps / SRE