Director of Engineering, Flex Compute
Leads the architecture, production launch, and team buildout for a safety-critical flexible-compute control system that manages AI infrastructure power in coordination with utilities. Requires extensive software engineering and team leadership experience, distributed orchestration expertise, and knowledge of data-center power systems and grid programs.
About the job
Responsibilities
- Own the curtailment orchestration and decision layer, including grid-signal ingestion, staged shedding, and per-SKU power capping without silently dropping paid workloads.
- Deliver a utility-validated pilot with fast ramp to setpoint, tight accuracy, high-fidelity telemetry, and a pre-commitment test harness.
- Build dynamic power management for oversubscription and workload-aware GPU power estimation validated against fleet telemetry.
- Design authenticated signal ingress, bounded blast radius, fail-safe defaults, and manual backstops for safety-critical operation.
- Own graceful ride-through of power-loss and grid-stress events, integrated with battery and generation backstops.
- Partner with Data Center Engineering and Energy teams on interconnection commitments, curtailment programs, and BESS/generation integration.
- Integrate with cloud control planes across Kubernetes and Slurm fleets.
- Make build-versus-leverage decisions across vendor power-management and grid-integration stacks.
- Hire and lead the engineering team from the ground up while keeping headcount sublinear to fleet growth.
Requirements
- 12+ years of software engineering experience, including 5+ years leading engineering teams.
- Experience taking systems from 0-to-1 through production scale.
- Deep experience with distributed control planes, orchestration, or fleet automation.
- Experience with safety-critical or physically actuating systems with bounded worst-case failure modes.
- Working knowledge of data-center power systems, including utility interconnection, switchgear, UPS, BESS, rack/PDU distribution, power telemetry, and GPU power management.
- Experience with energy markets or grid programs such as demand response, curtailable-load tariffs, or ISO/RTO market signals.
- Strong build-versus-buy judgment and comfort making decisions with incomplete data.
- Track record of hiring senior engineers and operating systems that must not fail silently.
Nice-to-haves
- GPU cluster operations or AI cloud infrastructure experience.
- Energy-sector experience in power generation, storage, utility software, or grid-scale controls.
- GPU power/performance modeling, DVFS, or productionizing research prototypes.
- Checkpointing, preemption, or power-aware scheduling for large training jobs.
Compensation & Benefits
- Compensation range of $285,000–$335,000 plus bonus.
- Restricted Stock Units included in all offers.
- Paid time off, paid holidays, and leave programs.
- Comprehensive health, dental, and vision insurance.
- Employer HSA contributions.
- Paid parental leave, life insurance, and short- and long-term disability coverage.
- Professional development and tuition reimbursement.
- Mental health and wellness support.
- Commuter benefits and cell phone stipend.
- 401(k) plan with company match up to 4% of salary.
- Volunteer time off, global travel insurance, emergency assistance, daily meal allowance, and location-specific programs.
Skills
Kubernetes, Slurm, Temporal, Distributed Systems, Control Planes, Fleet Automation, Grid Signals, Utility Interconnection, Switchgear, Ups, Bess, Gpu Power Management, Power Telemetry, Dvfs, Demand Response
Similar jobs
Engineering Management jobsLeads a team building foundational security services for Crusoe’s GPU cloud and infrastructure fleet, spanning identity, cryptography, runtime protection, access, vulnerability management, and architectural isolation. Requires 8+ years leading hands-on software or security engineering teams at large-scale infrastructure platforms.
Leads the technical direction, engineering quality, reliability, and hands-on architecture of a 30-person organization building payment, ledger, wallet, and settlement infrastructure. Requires 10+ years of production software experience, 4+ years leading engineers, and deep distributed-systems and regulated-finance expertise.
Leads architecture, development, and optimization of global Order-to-Cash revenue systems, integrating billing, ERP, sales, and financial platforms. The role requires extensive revenue systems experience, strong data engineering and API expertise, and the ability to lead ERP transformations and technical teams.
Leads the product and engineering strategy for AI-native enterprise applications across Finance, HR, and Legal, partnering with executives to transform workflows into intelligent products. Requires 15+ years in enterprise technology, product management, or software engineering and substantial multidisciplinary leadership experience.
Leads the organization responsible for deploying, sustaining, supporting, and improving integrated Hivemind software and hardware products in customer environments. Requires 15 years of technical operations or lifecycle leadership experience, systems engineering expertise, and experience building operational organizations.