Own the technical lifecycle and operational health of Runpod’s global high-density GPU fleet. Bridge hardware partners and engineering teams through hardware validation, network troubleshooting, AI-driven automation, incident response, and performance tuning for AI/ML workloads.
Salary not listedRemote3+ YOEDevOps / SRE
Site Reliability Engineer
RunpodUnited States
Site Reliability Engineer responsible for defining SLIs/SLOs, leading incident response, building observability (Prometheus/Grafana), automating toil, and driving production readiness for Runpod's AI cloud platform. Requires 5+ years SRE experience, strong Linux/distributed systems knowledge, and scripting skills.
150k – 200k/yrRemote5+ YOEDevOps / SRE
Search
Location
2 jobs
Job results
Datacenter Infrastructure Specialist
RunpodUnited States
Own the technical lifecycle and operational health of Runpod’s global high-density GPU fleet. Bridge hardware partners and engineering teams through hardware validation, network troubleshooting, AI-driven automation, incident response, and performance tuning for AI/ML workloads.
Salary not listedRemote3+ YOEDevOps / SRE
Site Reliability Engineer
RunpodUnited States
Site Reliability Engineer responsible for defining SLIs/SLOs, leading incident response, building observability (Prometheus/Grafana), automating toil, and driving production readiness for Runpod's AI cloud platform. Requires 5+ years SRE experience, strong Linux/distributed systems knowledge, and scripting skills.