Own fleet reliability for nation-scale AI data centers: set availability targets, perform RCA on major incidents, build failure data pipelines, and define data-driven maintenance strategies. Requires hands-on experience improving critical infrastructure availability using Weibull, Pareto, and FMEA.
398k – 557k/yr
On-site7+ YOEDevOps / SRE
About the role
Own fleet reliability engineering
Define availability targets, measure them honestly, and close the gap to targets.
Run root cause analysis on the fleet's worst incidents and drive corrective actions to completion across every site.
Build the failure data pipeline (facility and hardware) that turns incident history into engineering priorities.
Set the maintenance strategy (reliability-centered, condition-based) so the fleet spends effort where the failure data indicates.
Requirements
Owned reliability for critical infrastructure and moved the availability number (not just reported it).
Led root cause analyses that found the real cause, not the convenient one.
Work fluently with failure data: Weibull, Pareto, and FMEA are tools you actually use.
Get corrective actions closed across teams you do not manage.
Nice-to-haves
Data center or power generation reliability.
Liquid cooling systems.
CMMS analytics.
CRE or CMRP certification.
Skills
reliability engineeringRoot Cause Analysisweibull analysispareto analysisfmeafailure data pipelinemaintenance strategycmms
Leads reliability, scalability, and modernization of Pinterest's critical online systems including storage, caching, and real-time analytics at massive scale. Drives strategic vision, cross-functional initiatives like Kubernetes migration, requiring 12+ years in distributed systems and expertise in C++, Java, or Python.
285k – 500k/yr
Hybrid12+ YOEDevOps / SRE
Principal Engineer, CAPE
CrusoeSan Francisco, CA
Principal Engineer building Crusoe's self-driving Conductor platform for AI infrastructure. Own closed-loop autonomy, predictive failure detection, unified observability, energy-aware scheduling, and goodput optimization across tens of thousands of GPUs. Requires 10+ years in large-scale distributed systems, HPC/GPU infrastructure, and observability.
285k – 335k/yr
On-site10+ YOEDevOps / SRE
Principal Systems Engineer
BlacksmithNew York, NY
Principal Systems Engineer sets technical direction for core infrastructure, owns architecture for reliability and performance at scale, and mentors senior engineers. Requires deep expertise in virtualization, distributed storage like Ceph, and Linux kernel primitives.
280k – 380k/yr
On-siteDevOps / SRE
Principal Engineer, Compute Fleet Management
DatabricksBellevue, WA
Leads compute fleet management across AWS, Azure, and GCP, optimizing billions of resources for peak performance, 99.99% availability, and 60%+ utilization. Requires deep distributed systems expertise and cross-team leadership for mission-critical infrastructure.
264k – 322k/yr
On-siteDevOps / SRE
Principal Production Engineer
CrusoeSan Francisco, CA +1
Owns reliability, scalability, and observability of cloud infrastructure including compute, storage, and networking at massive scale. Drives SLOs, incident response, tooling, and mentors engineers; requires 15+ years experience with data centers and internet-scale operations.