Skip to content
CrusoeCrusoeSan Francisco, CA

Principal Engineer, CAPE

Principal Engineer building Crusoe's self-driving Conductor platform for AI infrastructure. Own closed-loop autonomy, predictive failure detection, unified observability, energy-aware scheduling, and goodput optimization across tens of thousands of GPUs. Requires 10+ years in large-scale distributed systems, HPC/GPU infrastructure, and observability.

285k – 335k/yr
On-site10+ YOEDevOps / SRE

About the role

Charter — The Problems You'll Own

  • Unified observability plane correlating GPU, networking (InfiniBand/RoCE), storage, orchestration, and workload signals for fast diagnosis and recovery.
  • Treat the fleet as one logical computer: tens of thousands of accelerators with one health model, scheduler, and source of truth.
  • Closed-loop autonomy: diagnose, decide, and remediate (drain, checkpoint, replace, resume) with no human in the loop, earning trust to act on live jobs.
  • Maximize goodput as the objective function, trading scheduling, placement, and maintenance decisions against it.
  • Predict failures hours ahead using ML on noisy hardware telemetry (GPU, NVLink, optics, thermal) and pre-emptively migrate work.
  • Straggler & silent-failure detection: isolate exact rank, GPU, and node from collective-operation signals.
  • Energy-aware compute: schedule, throttle, and place workloads against real-time energy availability, cost, and thermal headroom.
  • Digital twin of the fleet to simulate failures, scheduling policies, and remediation logic safely.
  • Agentic operations: an agent that proposes and executes fixes with guardrails and audit trails.
  • Zero-trust, fully auditable multi-tenancy with identity-scoped, policy-checked actions.
  • Self-qualifying hardware: automated burn-in for new/repaired nodes before taking load.

What You'll Bring

  • 10+ years building infrastructure-layer systems at scale (fleet management, distributed control planes, scheduler internals, or hardware lifecycle automation).
  • Deep experience with distributed systems design: consensus, state reconciliation, closed-loop automation, and autonomous decision-making on live production infrastructure.
  • Hands-on fluency with GPU/HPC infrastructure, including GPU health telemetry, NVLink/InfiniBand/RoCE fabrics, thermal and power behavior.
  • Track record designing and shipping large-scale observability or telemetry platforms correlating signals across compute, network, and storage.
  • Comfort operating in ambiguity and defining architecture/standards for 0→1 systems.
  • Strong software engineering fundamentals in at least one systems language (Go, Rust, C++, or similar) with judgment on build vs. adopt.
  • Experience applying ML/statistical methods to noisy operational telemetry (failure prediction, anomaly detection) is a strong plus.
  • Prior exposure to zero-trust or policy-based multi-tenancy architectures is a plus.

Benefits

  • Competitive compensation and equity packages including Restricted Stock Units.
  • Paid time off, paid holidays & leave of absence programs.
  • Comprehensive health, dental & vision insurance; employer contributions to HSA.
  • Paid parental leave; paid life insurance, short-term and long-term disability.
  • Professional development & tuition reimbursement.
  • Mental health & wellness support.
  • Commuter benefits (parking & transit); cell phone stipend.
  • 401(k) Retirement plan with company match up to 4% of salary.
  • Volunteer time off; global travel insurance & emergency assistance.
  • Daily meals allowance; additional perks & programs specific to location.

Skills

Distributed SystemsObservabilitytelemetrygpu infrastructurehpcInfiniBandrocenvlinkGoRustC++Machine LearningAnomaly DetectionZero Trustcontrol planes

Similar roles

DevOps / SRE jobs
Pinterest

Principal Engineer, Online Systems

PinterestSan Francisco, CA +1

Leads reliability, scalability, and modernization of Pinterest's critical online systems including storage, caching, and real-time analytics at massive scale. Drives strategic vision, cross-functional initiatives like Kubernetes migration, requiring 12+ years in distributed systems and expertise in C++, Java, or Python.

285k – 500k/yr
Hybrid12+ YOEDevOps / SRE
Blacksmith

Principal Systems Engineer

BlacksmithNew York, NY

Principal Systems Engineer sets technical direction for core infrastructure, owns architecture for reliability and performance at scale, and mentors senior engineers. Requires deep expertise in virtualization, distributed storage like Ceph, and Linux kernel primitives.

280k – 380k/yr
On-siteDevOps / SRE
Databricks

Principal Engineer, Compute Fleet Management

DatabricksBellevue, WA

Leads compute fleet management across AWS, Azure, and GCP, optimizing billions of resources for peak performance, 99.99% availability, and 60%+ utilization. Requires deep distributed systems expertise and cross-team leadership for mission-critical infrastructure.

264k – 322k/yr
On-siteDevOps / SRE
Crusoe

Principal Production Engineer

CrusoeSan Francisco, CA +1

Owns reliability, scalability, and observability of cloud infrastructure including compute, storage, and networking at massive scale. Drives SLOs, incident response, tooling, and mentors engineers; requires 15+ years experience with data centers and internet-scale operations.

261k – 326k/yr
On-site15+ YOEDevOps / SRE
Crusoe

Principal Systems Software Engineer

CrusoeSan Francisco, CA +1

Leads architecture of next-generation AI infrastructure, unifying BMaaS, IaaS, and CaaS with focus on high-performance I/O paths, kernel optimizations, and GPU workloads. Requires 12+ years hyperscale experience, deep Linux/virtualization expertise, and hardware-software co-design skills.

260k – 340k/yr
On-site12+ YOEDevOps / SRE