Skip to content
FluidstackFluidstackSan Francisco, CA

Network Production Engineering Lead

Lead the network production engineering team responsible for availability, performance, and automation of fabrics supporting 100k+ accelerator clusters at massive scale. Own SLOs, build remediation automation, and set operating models between design and site teams.

242k – 284k/yr
On-site7+ YOEDevOps / SRE

About the role

Role Scope

Lead the network production engineering team keeping fabrics for 100k+ accelerator clusters healthy. Own network availability and performance SLOs: link health, congestion, and failure response. Build automation for fabric operations: telemetry, anomaly detection, automated drain and repair. Set the operating model between design engineering and site operators so escalations flow clean in both directions.

Responsibilities

  • Operate at the scale of a nation, not a building. The fleet you run will draw more power than some countries, on the way to 10s to 100s of GWs.
  • Fly the plane while it's being built. Sites come online in pieces, and you keep the live ones running flawlessly while construction continues around them.
  • Write the playbook, don't inherit it. No prior operations org has run at this speed and scale, so the standards you set become the standard.
  • Lead network operations or production engineering for very large fabrics.
  • Automate network remediation at scale.
  • Read fabric telemetry and find the sick link before the training job does.
  • Run on-call programs teams didn't hate.

Requirements

  • Led network operations or production engineering for very large fabrics.
  • Automated network remediation at scale and trusted it enough to let it run.
  • Experience reading fabric telemetry to proactively identify issues.

Nice-to-Haves

  • AI or HPC fabrics.
  • InfiniBand and RoCE.
  • Network telemetry stacks.
  • Vendor TAC escalation management.

Skills

network operationsproduction engineeringnetwork automationtelemetryAnomaly DetectionInfiniBandrocehpc fabricsai fabrics

Similar roles

DevOps / SRE jobs
Convex

Senior Software Engineer, Infra/Systems

ConvexSan Francisco, CA

Convex is seeking a Senior Software Engineer to design, build, and maintain their global cloud infrastructure. This role involves working on core systems, improving performance and reliability, and owning architectural decisions.

240k+/yr
Hybrid6+ YOEDevOps / SRE
OpenAI

Software Engineer, Frontier Systems

OpenAISan Francisco, CA

Builds infrastructure to monitor, detect, remediate, and verify hardware health across global GPU/CPU clusters at hyperscale. Owns node lifecycle workflows and partners with teams to ensure compute reliability for AI training and inference. Requires 7+ years experience with Python, distributed systems, and operational tooling.

250k – 445k/yr
On-site7+ YOEDevOps / SRE
Decagon

Senior Software Engineer, Infrastructure

DecagonSan Francisco, CA +1

Senior Infrastructure Engineer designs, builds, and operates high-scale, low-latency production systems including networking, data, ML serving, and developer platforms. Requires 5+ years experience, strong SLO/observability skills, and expertise in areas like Kubernetes and multi-cloud.

250k – 330k/yr
On-site5+ YOEDevOps / SRE
Forward Networks

Site Reliability Engineer

Forward NetworksSanta Clara, CA

Build the Site Reliability Engineering function from the ground up at Forward, defining SLOs, building observability infrastructure, leading incident response, and embedding reliability into the SDLC for their complex SaaS platform. Requires 6+ years SRE/DevOps experience, strong networking and Kubernetes skills, and a track record maturing SRE practices.

230k – 250k/yr
On-site6+ YOEDevOps / SRE
Sphere

AI Agent Infrastructure Lead

SphereSan Francisco, CA

Leads development of internal AI agent infrastructure ("Goose") to boost velocity across engineering, ops, and other teams. Builds safe, autonomous agent workflows for codebase inspection, testing, and complex tasks with strong focus on safety and accuracy.

230k – 260k/yr
On-siteDevOps / SRE