Principal Engineer building Crusoe's self-driving Conductor platform for AI infrastructure. Own closed-loop autonomy, predictive failure detection, unified observability, energy-aware scheduling, and goodput optimization across tens of thousands of GPUs. Requires 10+ years in large-scale distributed systems, HPC/GPU infrastructure, and observability.
285k – 335k/yr
On-site10+ YOEDevOps / SRE
About the role
Charter — The Problems You'll Own
Unified observability plane correlating GPU, networking (InfiniBand/RoCE), storage, orchestration, and workload signals for fast diagnosis and recovery.
Treat the fleet as one logical computer: tens of thousands of accelerators with one health model, scheduler, and source of truth.
Closed-loop autonomy: diagnose, decide, and remediate (drain, checkpoint, replace, resume) with no human in the loop, earning trust to act on live jobs.
Maximize goodput as the objective function, trading scheduling, placement, and maintenance decisions against it.
Predict failures hours ahead using ML on noisy hardware telemetry (GPU, NVLink, optics, thermal) and pre-emptively migrate work.
Straggler & silent-failure detection: isolate exact rank, GPU, and node from collective-operation signals.
Energy-aware compute: schedule, throttle, and place workloads against real-time energy availability, cost, and thermal headroom.
Digital twin of the fleet to simulate failures, scheduling policies, and remediation logic safely.
Agentic operations: an agent that proposes and executes fixes with guardrails and audit trails.
Zero-trust, fully auditable multi-tenancy with identity-scoped, policy-checked actions.
Self-qualifying hardware: automated burn-in for new/repaired nodes before taking load.
What You'll Bring
10+ years building infrastructure-layer systems at scale (fleet management, distributed control planes, scheduler internals, or hardware lifecycle automation).
Deep experience with distributed systems design: consensus, state reconciliation, closed-loop automation, and autonomous decision-making on live production infrastructure.
Hands-on fluency with GPU/HPC infrastructure, including GPU health telemetry, NVLink/InfiniBand/RoCE fabrics, thermal and power behavior.
Track record designing and shipping large-scale observability or telemetry platforms correlating signals across compute, network, and storage.
Comfort operating in ambiguity and defining architecture/standards for 0→1 systems.
Strong software engineering fundamentals in at least one systems language (Go, Rust, C++, or similar) with judgment on build vs. adopt.
Experience applying ML/statistical methods to noisy operational telemetry (failure prediction, anomaly detection) is a strong plus.
Prior exposure to zero-trust or policy-based multi-tenancy architectures is a plus.
Benefits
Competitive compensation and equity packages including Restricted Stock Units.
Paid time off, paid holidays & leave of absence programs.
Comprehensive health, dental & vision insurance; employer contributions to HSA.
Paid parental leave; paid life insurance, short-term and long-term disability.
Leads reliability, scalability, and modernization of Pinterest's critical online systems including storage, caching, and real-time analytics at massive scale. Drives strategic vision, cross-functional initiatives like Kubernetes migration, requiring 12+ years in distributed systems and expertise in C++, Java, or Python.
285k – 500k/yr
Hybrid12+ YOEDevOps / SRE
Principal Systems Engineer
BlacksmithNew York, NY
Principal Systems Engineer sets technical direction for core infrastructure, owns architecture for reliability and performance at scale, and mentors senior engineers. Requires deep expertise in virtualization, distributed storage like Ceph, and Linux kernel primitives.
280k – 380k/yr
On-siteDevOps / SRE
Principal Engineer, Compute Fleet Management
DatabricksBellevue, WA
Leads compute fleet management across AWS, Azure, and GCP, optimizing billions of resources for peak performance, 99.99% availability, and 60%+ utilization. Requires deep distributed systems expertise and cross-team leadership for mission-critical infrastructure.
264k – 322k/yr
On-siteDevOps / SRE
Principal Production Engineer
CrusoeSan Francisco, CA +1
Owns reliability, scalability, and observability of cloud infrastructure including compute, storage, and networking at massive scale. Drives SLOs, incident response, tooling, and mentors engineers; requires 15+ years experience with data centers and internet-scale operations.
261k – 326k/yr
On-site15+ YOEDevOps / SRE
Principal Systems Software Engineer
CrusoeSan Francisco, CA +1
Leads architecture of next-generation AI infrastructure, unifying BMaaS, IaaS, and CaaS with focus on high-performance I/O paths, kernel optimizations, and GPU workloads. Requires 12+ years hyperscale experience, deep Linux/virtualization expertise, and hardware-software co-design skills.