Principal Engineer, Compute Fleet Management
Leads compute fleet management across AWS, Azure, and GCP, optimizing billions of resources for peak performance, 99.99% availability, and 60%+ utilization. Requires deep distributed systems expertise and cross-team leadership for mission-critical infrastructure.
About the job
Outcomes
- High Availability: Achieve and maintain 99.99% availability for all batch and serving workloads.
- Stellar Efficiency: Drive utilization to 60% or higher, balancing efficiency with tolerance for cloud failures.
- Best-in-Class Isolation: Architect and enforce strong security and performance isolation across diverse customer workloads.
Requirements
- Leading Transformative Projects: Take ownership of complex, cross-team, cross-layer, and multi-quarter strategic engineering initiatives from concept to execution.
- Distributed Systems Mastery: Deep, hands-on experience developing and operating high-scale distributed systems on at least one major public cloud.
- Influence Without Authority: Proven ability to drive consensus, establish technical direction, and lead large technical efforts across organizational boundaries.
- Execution Discipline: Exceptional strength in planning, tracking project progress, and managing complex cross-organizational dependencies.
The Edge: Highly Desirable Experience
- Experience managing and scaling a massive fleet of GPUs for AI/ML workloads.
- Experience with developing and operating large-scale distributed systems across all major clouds (AWS, Azure, and GCP).
Skills
Distributed Systems, AWS, Azure, GCP, Kubernetes, Fleet Management, GPU, Cloud Infrastructure, High Availability Systems, Resource Optimization
Similar jobs
DevOps / SRE jobsLeads Snowflake’s cloud infrastructure performance strategy by evaluating new hardware, building benchmark and validation systems, and translating performance data into pricing, capacity, and rollout decisions. Requires 12+ years in performance, systems, or infrastructure engineering and deep cloud hardware expertise.
This principal-level role owns operational excellence for a hyperscale AI data center network fleet, leading readiness, high-risk changes, audits, and incident resolution across sites. It requires extensive mission-critical network operations experience, routing and optical networking expertise, and 50–75% travel.
Build and operate AI-powered developer tools, internal MCP integrations, and platform capabilities across the engineering organization. The role requires strong coding and debugging skills, Kubernetes operations experience, and the ability to lead projects, improve developer experience, and mentor teammates.
As a Principal Operations Engineer, Mechanical, you will be the senior technical authority for mechanical and cooling infrastructure across hyperscale AI data centers. You will lead site assessments, drive operational readiness, review designs, and ensure precision execution of critical systems.
Own and scale secure cloud infrastructure, deployments, observability, compliance, and incident response for a hardware collaboration platform. The role requires substantial cloud or security engineering experience, AWS and Linux expertise, and the ability to lead cross-functional infrastructure initiatives.