# Datacenter Infrastructure Specialist

**Company:** [Runpod](https://hotfix.jobs/companies/runpod)
**Location:** Remote
**Role:** DevOps / SRE
**Experience:** 3+ years
**Skills:** nvidia software stack, Linux, Docker, rdma, InfiniBand, roce, Grafana, Prometheus, Datadog, Python, Go, Bash, hpc
**Posted:** 2026-07-29

> Own the technical lifecycle and operational health of Runpod’s global high-density GPU fleet. Bridge hardware partners and engineering teams through hardware validation, network troubleshooting, AI-driven automation, incident response, and performance tuning for AI/ML workloads.

## Job Description

## Responsibilities
- Assist in validating new hardware, ensuring partner deployments meet Runpod’s specifications for distributed AI/ML workloads.
- Monitor fleet health to identify performance degradation. Help audit downtime and provide technical data to protect customer SLAs.
- Work with LLMs and AI agents to automate network triage and generate dynamic runbooks for the fleet.
- Coordinate technical incident communications with clear updates, translating outages into actionable resolutions.
- Support the growth of infrastructure partners as the technical authority bridging hardware partners and internal engineering teams.
- Own the technical lifecycle and operational health of the high-density GPU fleet, including HPC systems engineering, advanced network troubleshooting, and process automation.

## Requirements
- 3–5 years of experience in infrastructure operations, systems reliability, or datacenter engineering.
- Strong proficiency in datacenter networking and performance troubleshooting.
- Hands-on experience with the NVIDIA Software Stack (driver installation, performance utilities) and understanding of multi-node performance tuning.
- Solid Linux system administration skills and experience with containerization (Docker).
- Comfortable with system-level troubleshooting and performance tuning at the kernel and hardware interface layers.
- Clear written and verbal communication skills to explain hardware or networking issues to technical partners and internal leadership.
- Detail-oriented and proactive in identifying potential failures before they impact customers.
- Willingness to participate in an on-call rotation as the global fleet scales.

## Nice-to-Haves
- Experience working in a fast-paced startup environment contributing to building operational workflows.
- Experience managing or optimizing bare-metal High-Performance Computing environments at massive scale.
- Experience with observability tools such as Grafana, Prometheus, or Datadog.
- Proficiency in Python, Go (Golang), or Bash to automate repetitive infrastructure tasks and interface with internal APIs.
- Exposure to RDMA, InfiniBand, or RoCE.

## Similar roles

- [Software Engineer, Compute Infrastructure](https://hotfix.jobs/jobs/e0a3b148-84c9-4527-a285-8b82c6473c63) - Glean - Mountain View, CA - $140k – $220k/yr
- [Software Engineer, Cloud Infrastructure](https://hotfix.jobs/jobs/14240d46-12c0-449f-a772-725a1df3ec0b) - Glean - Mountain View, CA - $200k – $270k/yr
- [Software Engineer, Cloud Deployment Infrastructure](https://hotfix.jobs/jobs/cb296000-b9c9-4d90-a725-05c15aff8b51) - Glean - San Francisco, CA - $200k – $270k/yr
- [Infrastructure engineer](https://hotfix.jobs/jobs/914e8c43-b09c-4819-96e1-24a8fb254fd7) - Writer - New York, NY - $140k – $274k/yr
- [Site Reliability Engineer](https://hotfix.jobs/jobs/f1fb6109-0b83-40f3-8972-d85f169649eb) - ConductorOne - Portland, ME - $180k – $250k/yr

**Apply:** https://hotfix.jobs/jobs/b85a0fe1-f849-48de-9fe6-702d3a47a159
**Canonical:** https://hotfix.jobs/jobs/b85a0fe1-f849-48de-9fe6-702d3a47a159