Skip to content
FluidstackFluidstackAustin, TX

Principal Operations Engineer, Reliability

Own fleet reliability for nation-scale AI data centers: set availability targets, perform RCA on major incidents, build failure data pipelines, and define data-driven maintenance strategies. Requires hands-on experience improving critical infrastructure availability using Weibull, Pareto, and FMEA.

398k – 557k/yr
On-site7+ YOEDevOps / SRE

About the role

Own fleet reliability engineering

  • Define availability targets, measure them honestly, and close the gap to targets.
  • Run root cause analysis on the fleet's worst incidents and drive corrective actions to completion across every site.
  • Build the failure data pipeline (facility and hardware) that turns incident history into engineering priorities.
  • Set the maintenance strategy (reliability-centered, condition-based) so the fleet spends effort where the failure data indicates.

Requirements

  • Owned reliability for critical infrastructure and moved the availability number (not just reported it).
  • Led root cause analyses that found the real cause, not the convenient one.
  • Work fluently with failure data: Weibull, Pareto, and FMEA are tools you actually use.
  • Get corrective actions closed across teams you do not manage.

Nice-to-haves

  • Data center or power generation reliability.
  • Liquid cooling systems.
  • CMMS analytics.
  • CRE or CMRP certification.

Skills

reliability engineeringRoot Cause Analysisweibull analysispareto analysisfmeafailure data pipelinemaintenance strategycmms

Similar roles

DevOps / SRE jobs
Pinterest

Principal Engineer, Online Systems

PinterestSan Francisco, CA +1

Leads reliability, scalability, and modernization of Pinterest's critical online systems including storage, caching, and real-time analytics at massive scale. Drives strategic vision, cross-functional initiatives like Kubernetes migration, requiring 12+ years in distributed systems and expertise in C++, Java, or Python.

285k – 500k/yr
Hybrid12+ YOEDevOps / SRE
Crusoe

Principal Engineer, CAPE

CrusoeSan Francisco, CA

Principal Engineer building Crusoe's self-driving Conductor platform for AI infrastructure. Own closed-loop autonomy, predictive failure detection, unified observability, energy-aware scheduling, and goodput optimization across tens of thousands of GPUs. Requires 10+ years in large-scale distributed systems, HPC/GPU infrastructure, and observability.

285k – 335k/yr
On-site10+ YOEDevOps / SRE
Blacksmith

Principal Systems Engineer

BlacksmithNew York, NY

Principal Systems Engineer sets technical direction for core infrastructure, owns architecture for reliability and performance at scale, and mentors senior engineers. Requires deep expertise in virtualization, distributed storage like Ceph, and Linux kernel primitives.

280k – 380k/yr
On-siteDevOps / SRE
Databricks

Principal Engineer, Compute Fleet Management

DatabricksBellevue, WA

Leads compute fleet management across AWS, Azure, and GCP, optimizing billions of resources for peak performance, 99.99% availability, and 60%+ utilization. Requires deep distributed systems expertise and cross-team leadership for mission-critical infrastructure.

264k – 322k/yr
On-siteDevOps / SRE
Crusoe

Principal Production Engineer

CrusoeSan Francisco, CA +1

Owns reliability, scalability, and observability of cloud infrastructure including compute, storage, and networking at massive scale. Drives SLOs, incident response, tooling, and mentors engineers; requires 15+ years experience with data centers and internet-scale operations.

261k – 326k/yr
On-site15+ YOEDevOps / SRE