Skip to content

Principal Operations Engineer, Network

This principal-level role owns operational excellence for a hyperscale AI data center network fleet, leading readiness, high-risk changes, audits, and incident resolution across sites. It requires extensive mission-critical network operations experience, routing and optical networking expertise, and 50–75% travel.

About the job

Responsibilities

  • Serve as the senior technical authority for the operational network fleet across hyperscale AI data centers, including switches, routers, copper and fiber cabling, and optics.
  • Lead site assessments and operational audits, and drive technical readiness for network operations teams before site activation.
  • Review network platforms and integration designs from an operational perspective, feeding lessons back into network engineering, deployment, and supply chain.
  • Author, approve, and execute high-risk network methods of procedure (MOPs) and change records in live production.
  • Lead fleet-wide root cause analyses of network disruptions through closure.
  • Hold OEMs, ODMs, and service vendors accountable to operational standards while maintaining effective relationships.
  • Coordinate across network operations, network engineering, compute operations, facilities, supply chain, and customer-facing teams.
  • Travel 50–75%.

Requirements

  • Career operating mission-critical network topologies at scale, with significant experience as the senior technical authority for a site, campus, or fleet.
  • Experience with IP network operations and optical networking.
  • Experience with TCP/IP and routing protocols including OSPF, IS-IS, BGP, and MPLS.
  • Experience with physical network infrastructure devices.
  • Experience authoring and executing high-risk MOPs and leading significant network-event root cause analyses through closure.
  • Experience managing OEMs, ODMs, and deployment partners to defined standards.
  • Strong technical writing skills for health assessments, RCAs, and design feedback.
  • Ability to teach and improve the technical readiness of surrounding teams.

Nice to Have

  • Experience with hyperscale or large HPC fleets supporting thousands of endpoints.
  • Linux experience.
  • Hardware management tooling.
  • Experience standing up new sites from handover through steady state.
  • Scripting for fleet-scale operations.

Compensation and Benefits

  • Total compensation of $258,000–$300,000, including salary and equity.
  • Retirement or pension plan in line with local norms.
  • Health, dental, and vision insurance.
  • Generous paid time off.
  • Equity may be provided as restricted stock units.

Skills

Ip Networking, Optical Networking, TCP/IP, Ospf, Is-Is, BGP, Mpls, Network Routing, Linux, Root Cause Analysis, Network Operations, Network Infrastructure, Scripting

Snowflake

Snowflake

Menlo Park, CA

Principal Software Engineer - Performance Engineering
$264k+/yrOn-site12+ YOEDevOps / SRE

Leads Snowflake’s cloud infrastructure performance strategy by evaluating new hardware, building benchmark and validation systems, and translating performance data into pricing, capacity, and rollout decisions. Requires 12+ years in performance, systems, or infrastructure engineering and deep cloud hardware expertise.

Cloudflare

Cloudflare

Atlanta, GA
Principal Systems Engineer, DevTools
$200k+/yrHybrid7+ YOEDevOps / SRE

Build and operate AI-powered developer tools, internal MCP integrations, and platform capabilities across the engineering organization. The role requires strong coding and debugging skills, Kubernetes operations experience, and the ability to lead projects, improve developer experience, and mentor teammates.

Fluidstack

Fluidstack

Remote

Principal Operations Engineer, Mechanical
$150k+/yrRemote10+ YOEDevOps / SRE

As a Principal Operations Engineer, Mechanical, you will be the senior technical authority for mechanical and cooling infrastructure across hyperscale AI data centers. You will lead site assessments, drive operational readiness, review designs, and ensure precision execution of critical systems.

AllSpice

AllSpice

Boston, MA
Principal / Staff / Senior Infrastructure Engineer
No salary listedHybrid8+ YOEDevOps / SRE

Own and scale secure cloud infrastructure, deployments, observability, compliance, and incident response for a hardware collaboration platform. The role requires substantial cloud or security engineering experience, AWS and Linux expertise, and the ability to lead cross-functional infrastructure initiatives.

Headway

Headway

San Francisco, CA
Staff Infrastructure Engineer
$265k+/yrRemote8+ YOEDevOps / SRE

Own the cloud platform, deployment architecture, container infrastructure, networking, autoscaling, cost controls, and Python runtime health for a high-scale healthcare technology platform. The role requires 8+ years in infrastructure, platform, or SRE work, deep AWS expertise, Terraform experience, and Staff-level cross-team influence.