Role Scope
Lead the network production engineering team keeping fabrics for 100k+ accelerator clusters healthy.
Own network availability and performance SLOs: link health, congestion, and failure response.
Build automation for fabric operations: telemetry, anomaly detection, automated drain and repair.
Set the operating model between design engineering and site operators so escalations flow clean in both directions.
Responsibilities
- Operate at the scale of a nation, not a building. The fleet you run will draw more power than some countries, on the way to 10s to 100s of GWs.
- Fly the plane while it's being built. Sites come online in pieces, and you keep the live ones running flawlessly while construction continues around them.
- Write the playbook, don't inherit it. No prior operations org has run at this speed and scale, so the standards you set become the standard.
- Lead network operations or production engineering for very large fabrics.
- Automate network remediation at scale.
- Read fabric telemetry and find the sick link before the training job does.
- Run on-call programs teams didn't hate.
Requirements
- Led network operations or production engineering for very large fabrics.
- Automated network remediation at scale and trusted it enough to let it run.
- Experience reading fabric telemetry to proactively identify issues.
Nice-to-Haves
- AI or HPC fabrics.
- InfiniBand and RoCE.
- Network telemetry stacks.
- Vendor TAC escalation management.