Skip to content

Principal Engineer, Inference Cloud

Principal Engineer leads Inference Cloud Platform, defining architecture for multi-region, high-QPS AI inference systems. Focuses on reliability, performance optimization, production code, and cross-team technical strategy. Requires 10+ years in distributed systems.

About the job

Responsibilities

  • Problem Definition & Prioritization: Identify the most important technical problems for the platform, make explicit tradeoff decisions.

  • Platform Direction: Set long-term technical direction including multi-region topology, failure domains, service boundaries, and system evolution.

  • Reliability & Performance: Architect active-active systems with rapid failover, graceful degradation (circuit breaking, backpressure, load shedding), SLOs; improve latency, throughput, capacity efficiency, resilience.

  • Code & Design Reviews: Contribute production code in critical paths, review designs and implementations, make architectural decisions including build-vs-buy.

  • Production Leadership: Lead on production issues, cross-system bottlenecks; drive observability, incident response, capacity planning, post-incident improvements.

  • Technical Strategy Beyond Your Team: Drive platform-wide decisions on reliability, API design, capacity planning, deployment strategy; translate product/business requirements into scalable designs.

  • Mentorship: Raise technical decision-making quality through design feedback, pairing, engineering standards.

Skills & Qualifications

  • 10+ years software engineering experience building/operating large-scale distributed systems or cloud infrastructure.

  • Deep expertise in distributed systems architecture: networking, compute orchestration, container platforms, multi-region services.

  • Track record of architectural decisions for highly available, latency-sensitive systems at scale.

  • Experience optimizing latency, throughput, efficiency in high-QPS systems (TTFT, tail-latency a plus).

  • Proficiency in Go, C++, or Python for production code.

  • Designing observability/reliability: metrics, logging, tracing, alerting, incident response, SLI/SLO/SLA.

  • Influence senior engineers/cross-functional partners via credibility, communication, judgment.

  • ML inference infrastructure, model serving, GPU workloads (plus).

Skills

Distributed Systems, Go, C++, Python, Kubernetes, Multi-Region Architecture, Observability, Slo/Sli, Latency Optimization, High-Qps Systems, Ml Inference, Circuit Breaking, Load Shedding, Incident Response, Capacity Planning

Fluidstack

Fluidstack

United States

Principal Operations Engineer, Network
$258k+/yrRemote7+ YOEDevOps / SRE

This principal-level role owns operational excellence for a hyperscale AI data center network fleet, leading readiness, high-risk changes, audits, and incident resolution across sites. It requires extensive mission-critical network operations experience, routing and optical networking expertise, and 50–75% travel.

Snowflake

Snowflake

Menlo Park, CA

Principal Software Engineer - Performance Engineering
$264k+/yrOn-site12+ YOEDevOps / SRE

Leads Snowflake’s cloud infrastructure performance strategy by evaluating new hardware, building benchmark and validation systems, and translating performance data into pricing, capacity, and rollout decisions. Requires 12+ years in performance, systems, or infrastructure engineering and deep cloud hardware expertise.

AllSpice

AllSpice

Boston, MA
Principal / Staff / Senior Infrastructure Engineer
No salary listedHybrid8+ YOEDevOps / SRE

Own and scale secure cloud infrastructure, deployments, observability, compliance, and incident response for a hardware collaboration platform. The role requires substantial cloud or security engineering experience, AWS and Linux expertise, and the ability to lead cross-functional infrastructure initiatives.

Fluidstack

Fluidstack

Remote

Principal Operations Engineer, Mechanical
$150k+/yrRemote10+ YOEDevOps / SRE

As a Principal Operations Engineer, Mechanical, you will be the senior technical authority for mechanical and cooling infrastructure across hyperscale AI data centers. You will lead site assessments, drive operational readiness, review designs, and ensure precision execution of critical systems.

Cloudflare

Cloudflare

Atlanta, GA
Principal Systems Engineer, DevTools
$200k+/yrHybrid7+ YOEDevOps / SRE

Build and operate AI-powered developer tools, internal MCP integrations, and platform capabilities across the engineering organization. The role requires strong coding and debugging skills, Kubernetes operations experience, and the ability to lead projects, improve developer experience, and mentor teammates.