# Principal Engineer, CAPE

**Company:** [Crusoe](https://hotfix.jobs/companies/crusoe)
**Location:** San Francisco, CA
**Role:** DevOps / SRE
**Salary:** $285k – $335k/yr
**Experience:** 10+ years
**Skills:** Distributed Systems, Observability, telemetry, gpu infrastructure, hpc, InfiniBand, roce, nvlink, Go, Rust, C++, Machine Learning, Anomaly Detection, Zero Trust, control planes
**Posted:** 2026-07-21

> Principal Engineer building Crusoe's self-driving Conductor platform for AI infrastructure. Own closed-loop autonomy, predictive failure detection, unified observability, energy-aware scheduling, and goodput optimization across tens of thousands of GPUs. Requires 10+ years in large-scale distributed systems, HPC/GPU infrastructure, and observability.

## Job Description

## Charter — The Problems You'll Own
- Unified observability plane correlating GPU, networking (InfiniBand/RoCE), storage, orchestration, and workload signals for fast diagnosis and recovery.
- Treat the fleet as one logical computer: tens of thousands of accelerators with one health model, scheduler, and source of truth.
- Closed-loop autonomy: diagnose, decide, and remediate (drain, checkpoint, replace, resume) with no human in the loop, earning trust to act on live jobs.
- Maximize goodput as the objective function, trading scheduling, placement, and maintenance decisions against it.
- Predict failures hours ahead using ML on noisy hardware telemetry (GPU, NVLink, optics, thermal) and pre-emptively migrate work.
- Straggler & silent-failure detection: isolate exact rank, GPU, and node from collective-operation signals.
- Energy-aware compute: schedule, throttle, and place workloads against real-time energy availability, cost, and thermal headroom.
- Digital twin of the fleet to simulate failures, scheduling policies, and remediation logic safely.
- Agentic operations: an agent that proposes and executes fixes with guardrails and audit trails.
- Zero-trust, fully auditable multi-tenancy with identity-scoped, policy-checked actions.
- Self-qualifying hardware: automated burn-in for new/repaired nodes before taking load.

## What You'll Bring
- 10+ years building infrastructure-layer systems at scale (fleet management, distributed control planes, scheduler internals, or hardware lifecycle automation).
- Deep experience with distributed systems design: consensus, state reconciliation, closed-loop automation, and autonomous decision-making on live production infrastructure.
- Hands-on fluency with GPU/HPC infrastructure, including GPU health telemetry, NVLink/InfiniBand/RoCE fabrics, thermal and power behavior.
- Track record designing and shipping large-scale observability or telemetry platforms correlating signals across compute, network, and storage.
- Comfort operating in ambiguity and defining architecture/standards for 0→1 systems.
- Strong software engineering fundamentals in at least one systems language (Go, Rust, C++, or similar) with judgment on build vs. adopt.
- Experience applying ML/statistical methods to noisy operational telemetry (failure prediction, anomaly detection) is a strong plus.
- Prior exposure to zero-trust or policy-based multi-tenancy architectures is a plus.

## Benefits
- Competitive compensation and equity packages including Restricted Stock Units.
- Paid time off, paid holidays & leave of absence programs.
- Comprehensive health, dental & vision insurance; employer contributions to HSA.
- Paid parental leave; paid life insurance, short-term and long-term disability.
- Professional development & tuition reimbursement.
- Mental health & wellness support.
- Commuter benefits (parking & transit); cell phone stipend.
- 401(k) Retirement plan with company match up to 4% of salary.
- Volunteer time off; global travel insurance & emergency assistance.
- Daily meals allowance; additional perks & programs specific to location.

## Similar roles

- [Principal Engineer, Online Systems](https://hotfix.jobs/jobs/055c1895-ab91-4500-9b34-b24e8c0036ca) - Pinterest - San Francisco, CA - $285k – $500k/yr
- [Principal Systems Engineer](https://hotfix.jobs/jobs/a978d66d-da58-4c11-9dae-d9fc07c6a3c8) - Blacksmith - New York, NY - $280k – $380k/yr
- [Principal Engineer, Compute Fleet Management](https://hotfix.jobs/jobs/ee6ebd92-809c-4ee8-a361-9e58b7593a17) - Databricks - Bellevue, WA - $264k – $322k/yr
- [Principal Production Engineer](https://hotfix.jobs/jobs/2025d56d-ab2d-4a71-a0e8-ad577d1029c6) - Crusoe - San Francisco, CA - $261k – $326k/yr
- [Principal Systems Software Engineer](https://hotfix.jobs/jobs/423c3fe7-aab5-4133-8cee-8ecfa36ad082) - Crusoe - San Francisco, CA - $260k – $340k/yr

**Apply:** https://hotfix.jobs/jobs/d718f9cb-5e2e-44b4-a653-81025cfceffb
**Canonical:** https://hotfix.jobs/jobs/d718f9cb-5e2e-44b4-a653-81025cfceffb