Skip to content
DatabricksDatabricks

Principal Engineer, Compute Fleet Management

Leads compute fleet management across AWS, Azure, and GCP, optimizing billions of resources for peak performance, 99.99% availability, and 60%+ utilization. Requires deep distributed systems expertise and cross-team leadership for mission-critical infrastructure.

About the job

Outcomes

  • High Availability: Achieve and maintain 99.99% availability for all batch and serving workloads.
  • Stellar Efficiency: Drive utilization to 60% or higher, balancing efficiency with tolerance for cloud failures.
  • Best-in-Class Isolation: Architect and enforce strong security and performance isolation across diverse customer workloads.

Requirements

  • Leading Transformative Projects: Take ownership of complex, cross-team, cross-layer, and multi-quarter strategic engineering initiatives from concept to execution.
  • Distributed Systems Mastery: Deep, hands-on experience developing and operating high-scale distributed systems on at least one major public cloud.
  • Influence Without Authority: Proven ability to drive consensus, establish technical direction, and lead large technical efforts across organizational boundaries.
  • Execution Discipline: Exceptional strength in planning, tracking project progress, and managing complex cross-organizational dependencies.

The Edge: Highly Desirable Experience

  • Experience managing and scaling a massive fleet of GPUs for AI/ML workloads.
  • Experience with developing and operating large-scale distributed systems across all major clouds (AWS, Azure, and GCP).

Skills

Distributed Systems, AWS, Azure, GCP, Kubernetes, Fleet Management, GPU, Cloud Infrastructure, High Availability Systems, Resource Optimization

Snowflake

Snowflake

Menlo Park, CA

Principal Software Engineer - Performance Engineering
$264k+/yrOn-site12+ YOEDevOps / SRE

Leads Snowflake’s cloud infrastructure performance strategy by evaluating new hardware, building benchmark and validation systems, and translating performance data into pricing, capacity, and rollout decisions. Requires 12+ years in performance, systems, or infrastructure engineering and deep cloud hardware expertise.

Fluidstack

Fluidstack

United States

Principal Operations Engineer, Network
$258k+/yrRemote7+ YOEDevOps / SRE

This principal-level role owns operational excellence for a hyperscale AI data center network fleet, leading readiness, high-risk changes, audits, and incident resolution across sites. It requires extensive mission-critical network operations experience, routing and optical networking expertise, and 50–75% travel.

Cloudflare

Cloudflare

Atlanta, GA
Principal Systems Engineer, DevTools
$200k+/yrHybrid7+ YOEDevOps / SRE

Build and operate AI-powered developer tools, internal MCP integrations, and platform capabilities across the engineering organization. The role requires strong coding and debugging skills, Kubernetes operations experience, and the ability to lead projects, improve developer experience, and mentor teammates.

Fluidstack

Fluidstack

Remote

Principal Operations Engineer, Mechanical
$150k+/yrRemote10+ YOEDevOps / SRE

As a Principal Operations Engineer, Mechanical, you will be the senior technical authority for mechanical and cooling infrastructure across hyperscale AI data centers. You will lead site assessments, drive operational readiness, review designs, and ensure precision execution of critical systems.

AllSpice

AllSpice

Boston, MA
Principal / Staff / Senior Infrastructure Engineer
No salary listedHybrid8+ YOEDevOps / SRE

Own and scale secure cloud infrastructure, deployments, observability, compliance, and incident response for a hardware collaboration platform. The role requires substantial cloud or security engineering experience, AWS and Linux expertise, and the ability to lead cross-functional infrastructure initiatives.