# Network Production Engineering Lead

**Company:** [Fluidstack](https://hotfix.jobs/companies/fluidstack)
**Location:** San Francisco, CA
**Role:** DevOps / SRE
**Salary:** $242k – $284k/yr
**Experience:** 7+ years
**Skills:** network operations, production engineering, network automation, telemetry, Anomaly Detection, InfiniBand, roce, hpc fabrics, ai fabrics
**Posted:** 2026-07-21

> Lead the network production engineering team responsible for availability, performance, and automation of fabrics supporting 100k+ accelerator clusters at massive scale. Own SLOs, build remediation automation, and set operating models between design and site teams.

## Job Description

## Role Scope
Lead the network production engineering team keeping fabrics for 100k+ accelerator clusters healthy.
Own network availability and performance SLOs: link health, congestion, and failure response.
Build automation for fabric operations: telemetry, anomaly detection, automated drain and repair.
Set the operating model between design engineering and site operators so escalations flow clean in both directions.

## Responsibilities
- Operate at the scale of a nation, not a building. The fleet you run will draw more power than some countries, on the way to 10s to 100s of GWs.
- Fly the plane while it's being built. Sites come online in pieces, and you keep the live ones running flawlessly while construction continues around them.
- Write the playbook, don't inherit it. No prior operations org has run at this speed and scale, so the standards you set become the standard.
- Lead network operations or production engineering for very large fabrics.
- Automate network remediation at scale.
- Read fabric telemetry and find the sick link before the training job does.
- Run on-call programs teams didn't hate.

## Requirements
- Led network operations or production engineering for very large fabrics.
- Automated network remediation at scale and trusted it enough to let it run.
- Experience reading fabric telemetry to proactively identify issues.

## Nice-to-Haves
- AI or HPC fabrics.
- InfiniBand and RoCE.
- Network telemetry stacks.
- Vendor TAC escalation management.

## Similar roles

- [Senior Software Engineer, Infra/Systems](https://hotfix.jobs/jobs/17f71b1c-4c11-4ef1-bdea-b4ff123df84d) - Convex - San Francisco, CA - From $240k/yr
- [Software Engineer, Frontier Systems](https://hotfix.jobs/jobs/5a732c74-cdbe-4fc9-9ba3-9be015a80279) - OpenAI - San Francisco, CA - $250k – $445k/yr
- [Senior Software Engineer, Infrastructure](https://hotfix.jobs/jobs/def51027-c615-4950-acba-e46662dd6b9a) - Decagon - San Francisco, CA - $250k – $330k/yr
- [Site Reliability Engineer](https://hotfix.jobs/jobs/ab64f10b-7c19-4c49-b827-0e3559c514c9) - Forward Networks - Santa Clara, CA - $230k – $250k/yr
- [AI Agent Infrastructure Lead](https://hotfix.jobs/jobs/56a036e4-396a-4890-a059-6e706796fe04) - Sphere - San Francisco, CA - $230k – $260k/yr

**Apply:** https://hotfix.jobs/jobs/51065552-0f77-47bc-a8f2-edf2c84caecc
**Canonical:** https://hotfix.jobs/jobs/51065552-0f77-47bc-a8f2-edf2c84caecc