Skip to content
FluidstackFluidstackSan Francsisco, CA

Software Engineer, Networking

Build and own end-to-end network fleet health, monitoring, debugging tooling, and automated repair pipelines for one of the world's largest AI datacenter networks. Requires systems thinking, on-call ownership, and fluency with AI coding tools plus Go/Python network automation experience.

208k – 269k/yr
On-site5+ YOEDevOps / SRE

About the role

Responsibilities

  • Own network fleet health end to end: define realtime monitoring requirements, build alerting lifecycle, and ship dashboards for network state across all sites.
  • Build active debugging tooling including link diagnostics, remote command execution across the fleet, and repair visualization.
  • Turn repair into a pipeline: build automation from detection through parts management, return to service, ticket integration, repair lifecycle pipelines, transceiver and optics tracking.
  • Own network qualification and validation: build frameworks that gate new sites and hardware into production; define what a healthy network looks like.
  • Own end-to-end reliability, scalability, and operation of the network at-scale through aggressive automation, tooling, and incident discipline.

Requirements

  • Treat toil as a bug and build tools to eliminate manual work (e.g., instead of manual SSH and commands).
  • Think in systems: understand how faults (transceiver, route, power) propagate and build tooling to distinguish them.
  • Move toward ambiguity, build the map, and explain it to others.
  • Learn at a steep slope and reach competence in unfamiliar domains quickly.
  • Comfortable carrying a pager, running incidents, writing postmortems, and fixing systemic causes.
  • Fluent with AI tooling: LLM APIs, MCP servers, agentic frameworks; drive Claude Code, Cursor, or similar daily.
  • Shipped production network tooling or automation that other teams depend on; comfortable in any language using AI coding tools.
  • Developed automation tools in Go and Python.
  • Experienced in link diagnostics, optical networks, and network monitoring (gNMI, gRPC, NETCONF, SONiC).

Nice-to-Haves

  • RMA and repair lifecycle automation.
  • Large-scale datacenter fabric (BGP, ECMP, spine-leaf).
  • Out-of-band network management.

Skills

GoPythongnmigRPCnetconfsonicLLM APIsai coding toolsnetwork monitoringlink diagnosticsoptical networksBGPecmp

Similar roles

DevOps / SRE jobs
Fluidstack

Software Engineer, Compute

FluidstackSan Francisco, CA +3

Build and own automation, observability, and repair pipelines for one of the world's largest GPU compute fleets. Requires hardware intuition at the firmware/silicon level, on-call ownership, and fluency with AI coding tools to eliminate toil at hyperscale.

208k – 269k/yr
On-site5+ YOEDevOps / SRE
Fluidstack

Site Reliability Engineer, Compute

FluidstackSan Francisco, CA +3

Own end-to-end health, reliability, and automation of a massive GPU compute fleet for AI infrastructure. Build metrics, alerting, repair pipelines, GPU qualification platforms, and low-level BMC/Redfish tooling while driving incidents and using AI coding tools daily.

208k – 269k/yr
On-site5+ YOEDevOps / SRE
Fluidstack

Network Software Engineer

FluidstackSan Francisco, CA +3

Build and own end-to-end network monitoring, debugging tooling, repair automation pipelines, and qualification frameworks for Fluidstack's massive AI datacenter fleet. Requires systems thinking, on-call ownership, Go/Python automation experience, network protocols expertise, and daily use of AI coding tools.

208k – 269k/yr
On-site5+ YOEDevOps / SRE
Fluidstack

Network Automation Engineer

FluidstackSan Francisco, CA +3

Own end-to-end network fleet health, monitoring, debugging tooling, and automated repair pipelines for massive AI datacenter infrastructure. Requires systems thinking, building automation in Go/Python, experience with optical networks and protocols like gNMI/gRPC/NETCONF/SONiC, and daily use of AI coding tools.

208k – 269k/yr
On-site5+ YOEDevOps / SRE
Fluidstack

Distributed Systems Engineer

FluidstackSan Francisco, CA +3

Build and own the observability platform, production control plane, and fleet state as source of truth for a hyperscale GPU fleet powering AI infrastructure. Requires shipping scalable production services, on-call ownership, and comfort with AI coding tools; distributed systems and observability experience preferred.

208k – 269k/yr
On-site5+ YOEDevOps / SRE