Build and own end-to-end network fleet health, monitoring, debugging tooling, and automated repair pipelines for one of the world's largest AI datacenter networks. Requires systems thinking, on-call ownership, and fluency with AI coding tools plus Go/Python network automation experience.
208k – 269k/yr
On-site5+ YOEDevOps / SRE
About the role
Responsibilities
Own network fleet health end to end: define realtime monitoring requirements, build alerting lifecycle, and ship dashboards for network state across all sites.
Build active debugging tooling including link diagnostics, remote command execution across the fleet, and repair visualization.
Turn repair into a pipeline: build automation from detection through parts management, return to service, ticket integration, repair lifecycle pipelines, transceiver and optics tracking.
Own network qualification and validation: build frameworks that gate new sites and hardware into production; define what a healthy network looks like.
Own end-to-end reliability, scalability, and operation of the network at-scale through aggressive automation, tooling, and incident discipline.
Requirements
Treat toil as a bug and build tools to eliminate manual work (e.g., instead of manual SSH and commands).
Think in systems: understand how faults (transceiver, route, power) propagate and build tooling to distinguish them.
Move toward ambiguity, build the map, and explain it to others.
Learn at a steep slope and reach competence in unfamiliar domains quickly.
Comfortable carrying a pager, running incidents, writing postmortems, and fixing systemic causes.
Fluent with AI tooling: LLM APIs, MCP servers, agentic frameworks; drive Claude Code, Cursor, or similar daily.
Shipped production network tooling or automation that other teams depend on; comfortable in any language using AI coding tools.
Developed automation tools in Go and Python.
Experienced in link diagnostics, optical networks, and network monitoring (gNMI, gRPC, NETCONF, SONiC).
Build and own automation, observability, and repair pipelines for one of the world's largest GPU compute fleets. Requires hardware intuition at the firmware/silicon level, on-call ownership, and fluency with AI coding tools to eliminate toil at hyperscale.
208k – 269k/yr
On-site5+ YOEDevOps / SRE
Site Reliability Engineer, Compute
FluidstackSan Francisco, CA +3
Own end-to-end health, reliability, and automation of a massive GPU compute fleet for AI infrastructure. Build metrics, alerting, repair pipelines, GPU qualification platforms, and low-level BMC/Redfish tooling while driving incidents and using AI coding tools daily.
208k – 269k/yr
On-site5+ YOEDevOps / SRE
Network Software Engineer
FluidstackSan Francisco, CA +3
Build and own end-to-end network monitoring, debugging tooling, repair automation pipelines, and qualification frameworks for Fluidstack's massive AI datacenter fleet. Requires systems thinking, on-call ownership, Go/Python automation experience, network protocols expertise, and daily use of AI coding tools.
208k – 269k/yr
On-site5+ YOEDevOps / SRE
Network Automation Engineer
FluidstackSan Francisco, CA +3
Own end-to-end network fleet health, monitoring, debugging tooling, and automated repair pipelines for massive AI datacenter infrastructure. Requires systems thinking, building automation in Go/Python, experience with optical networks and protocols like gNMI/gRPC/NETCONF/SONiC, and daily use of AI coding tools.
208k – 269k/yr
On-site5+ YOEDevOps / SRE
Distributed Systems Engineer
FluidstackSan Francisco, CA +3
Build and own the observability platform, production control plane, and fleet state as source of truth for a hyperscale GPU fleet powering AI infrastructure. Requires shipping scalable production services, on-call ownership, and comfort with AI coding tools; distributed systems and observability experience preferred.