Principal Engineer, CAPE
Principal Engineer building Crusoe's self-driving Conductor platform for AI infrastructure. Own closed-loop autonomy, predictive failure detection, unified observability, energy-aware scheduling, and goodput optimization across tens of thousands of GPUs. Requires 10+ years in large-scale distributed systems, HPC/GPU infrastructure, and observability.
About the job
Charter — The Problems You'll Own
- Unified observability plane correlating GPU, networking (InfiniBand/RoCE), storage, orchestration, and workload signals for fast diagnosis and recovery.
- Treat the fleet as one logical computer: tens of thousands of accelerators with one health model, scheduler, and source of truth.
- Closed-loop autonomy: diagnose, decide, and remediate (drain, checkpoint, replace, resume) with no human in the loop, earning trust to act on live jobs.
- Maximize goodput as the objective function, trading scheduling, placement, and maintenance decisions against it.
- Predict failures hours ahead using ML on noisy hardware telemetry (GPU, NVLink, optics, thermal) and pre-emptively migrate work.
- Straggler & silent-failure detection: isolate exact rank, GPU, and node from collective-operation signals.
- Energy-aware compute: schedule, throttle, and place workloads against real-time energy availability, cost, and thermal headroom.
- Digital twin of the fleet to simulate failures, scheduling policies, and remediation logic safely.
- Agentic operations: an agent that proposes and executes fixes with guardrails and audit trails.
- Zero-trust, fully auditable multi-tenancy with identity-scoped, policy-checked actions.
- Self-qualifying hardware: automated burn-in for new/repaired nodes before taking load.
What You'll Bring
- 10+ years building infrastructure-layer systems at scale (fleet management, distributed control planes, scheduler internals, or hardware lifecycle automation).
- Deep experience with distributed systems design: consensus, state reconciliation, closed-loop automation, and autonomous decision-making on live production infrastructure.
- Hands-on fluency with GPU/HPC infrastructure, including GPU health telemetry, NVLink/InfiniBand/RoCE fabrics, thermal and power behavior.
- Track record designing and shipping large-scale observability or telemetry platforms correlating signals across compute, network, and storage.
- Comfort operating in ambiguity and defining architecture/standards for 0→1 systems.
- Strong software engineering fundamentals in at least one systems language (Go, Rust, C++, or similar) with judgment on build vs. adopt.
- Experience applying ML/statistical methods to noisy operational telemetry (failure prediction, anomaly detection) is a strong plus.
- Prior exposure to zero-trust or policy-based multi-tenancy architectures is a plus.
Benefits
- Competitive compensation and equity packages including Restricted Stock Units.
- Paid time off, paid holidays & leave of absence programs.
- Comprehensive health, dental & vision insurance; employer contributions to HSA.
- Paid parental leave; paid life insurance, short-term and long-term disability.
- Professional development & tuition reimbursement.
- Mental health & wellness support.
- Commuter benefits (parking & transit); cell phone stipend.
- 401(k) Retirement plan with company match up to 4% of salary.
- Volunteer time off; global travel insurance & emergency assistance.
- Daily meals allowance; additional perks & programs specific to location.
Skills
Distributed Systems, Observability, Telemetry, Gpu Infrastructure, Hpc, InfiniBand, Roce, Nvlink, Go, Rust, C++, Machine Learning, Anomaly Detection, Zero Trust, Control Planes
Similar jobs
DevOps / SRE jobsLeads Snowflake’s cloud infrastructure performance strategy by evaluating new hardware, building benchmark and validation systems, and translating performance data into pricing, capacity, and rollout decisions. Requires 12+ years in performance, systems, or infrastructure engineering and deep cloud hardware expertise.
This principal-level role owns operational excellence for a hyperscale AI data center network fleet, leading readiness, high-risk changes, audits, and incident resolution across sites. It requires extensive mission-critical network operations experience, routing and optical networking expertise, and 50–75% travel.
Build and operate AI-powered developer tools, internal MCP integrations, and platform capabilities across the engineering organization. The role requires strong coding and debugging skills, Kubernetes operations experience, and the ability to lead projects, improve developer experience, and mentor teammates.
As a Principal Operations Engineer, Mechanical, you will be the senior technical authority for mechanical and cooling infrastructure across hyperscale AI data centers. You will lead site assessments, drive operational readiness, review designs, and ensure precision execution of critical systems.
Own and scale secure cloud infrastructure, deployments, observability, compliance, and incident response for a hardware collaboration platform. The role requires substantial cloud or security engineering experience, AWS and Linux expertise, and the ability to lead cross-functional infrastructure initiatives.