Skip to content
Lightning AILightning AINew York, NY

Senior Network Engineer

Senior Network Engineer responsible for designing, deploying, and optimizing large-scale NVIDIA InfiniBand fabrics and UFM for AI/ML GPU clusters. Requires 10+ years data center networking experience with deep expertise in InfiniBand, spine-leaf architectures, automation, and HPC environments.

170k – 210k/yr
On-site10+ YOEDevOps / SRE

About the role

What You’ll Do

  • Design, deploy, and maintain large-scale NVIDIA InfiniBand fabrics supporting AI/ML GPU clusters.
  • Deploy and administer NVIDIA Unified Fabric Manager (UFM) Enterprise for monitoring, provisioning, telemetry, and fabric health.
  • Configure and optimize NVIDIA Quantum and Quantum-2 InfiniBand switches.
  • Troubleshoot fabric performance issues impacting NCCL, MPI, GPUDirect RDMA, and AI training jobs.
  • Implement and validate fat-tree, Dragonfly+, Clos, and spine-leaf network architectures.
  • Perform firmware lifecycle management for InfiniBand switches, adapters (HCAs), and UFM infrastructure.
  • Optimize congestion control, adaptive routing, QoS, and traffic engineering for high-performance GPU communication.
  • Work closely with AI platform, GPU infrastructure, storage, and systems engineering teams to deploy scalable AI Factory environments.
  • Automate network provisioning using Python, Ansible, Git, REST APIs, and Infrastructure-as-Code methodologies.
  • Monitor network health using UFM telemetry, Prometheus, Grafana, and other observability platforms.
  • Support high availability, maintenance windows, incident response, root cause analysis, and capacity planning.
  • Participate in architecture reviews and define networking standards for AI infrastructure.

Required Qualifications

  • 7+ years of data center networking experience.
  • 3+ years supporting NVIDIA InfiniBand environments.
  • Hands-on experience with NVIDIA UFM Enterprise.
  • Experience deploying and operating Quantum and Quantum-2 InfiniBand switches.
  • Strong understanding of InfiniBand Architecture, Subnet Manager (SM), Adaptive Routing, Congestion Control, Partition Keys (PKeys), LIDs, Queue Pairs (QP), Virtual Lanes (VL), Service Levels (SL).
  • Strong Linux administration experience (Ubuntu).
  • Experience with automation using Python and Ansible.
  • Deep understanding of Layer 2 and Layer 3 networking.
  • Experience with BGP, EVPN, VXLAN, and modern spine-leaf architectures.
  • Experience with packet captures and troubleshooting using tcpdump, Wireshark, and ibdiagnet tools.
  • Excellent troubleshooting and communication skills.

What You Bring

  • 10+ years of experience in large-scale data center networking.
  • Deep expertise in spine-leaf architectures and L3 fabrics.
  • Strong experience with BGP, EVPN, VXLAN.
  • Experience operating high-performance computing (HPC) or GPU-dense environments.
  • Experience designing networks for hyperscalers, neoclouds, or high-scale SaaS infrastructure.
  • Strong automation background (Python, Ansible, Terraform, or similar).
  • Experience with network observability tooling and telemetry pipelines.
  • Proven ability to design systems that scale to thousands of nodes.
  • Strong documentation and architectural communication skills.

Nice to Have

  • Experience with Netris and Terraform.
  • Experience with multi-region backbone design.
  • Exposure to bare-metal provisioning systems.
  • Experience working in high-growth infrastructure startups.

Benefits and Perks

  • Comprehensive medical, dental and vision coverage (U.S.); Private medical and dental insurance (U.K.)
  • Retirement and financial wellness support (U.S.); Pension contribution (U.K.)
  • Generous paid time off, plus holidays
  • Paid parental leave
  • Professional development support
  • Wellness and work-from-home stipends
  • Flexible work environment

Skills

InfiniBandufm enterprisenvidia quantumnvidia quantum-2subnet manageradaptive routingcongestion controlPythonAnsibleLinuxBGPevpnvxlanPrometheusGrafana

Similar roles

DevOps / SRE jobs
Crusoe

Senior Performance Engineer

CrusoeSan Francisco, CA

Senior Performance Engineer responsible for Linux kernel optimization, system benchmarking, and low-level performance tuning to enhance Crusoe's AI cloud infrastructure. Requires deep Linux kernel expertise, proficiency in Go/C/C++, and hands-on experience with performance optimization in complex environments.

170k – 205k/yr
On-site5+ YOEDevOps / SRE
Illumio

Sr. Site Reliability Engineer

IllumioSunnyvale, CA

Senior Site Reliability Engineer responsible for monitoring, incident response, and optimizing the reliability, scalability, and performance of Illumio's AWS and Azure cloud infrastructure and SaaS services. Requires 5+ years SRE experience with strong cloud platform expertise.

170k – 196k/yr
On-site5+ YOEDevOps / SRE
Render

Software Engineer, Compute Infrastructure

RenderSan Francisco, CA

Build and own core compute infrastructure for Render's cloud platform, including Kubernetes clusters on hyperscalers and bare metal. Design, scale, debug, and optimize large-scale orchestration, scheduling, and distributed systems with deep Kubernetes and systems expertise.

170k – 290k/yr
Remote7+ YOEDevOps / SRE
Skydio

Senior Software Engineer, Developer Productivity Cloud Infrastructure

SkydioSan Mateo, CA

Senior engineer focused on developer productivity and cloud infrastructure. Designs scalable internal tools, re-architects build systems, and improves CI/CD workflows using Terraform, Go/Python/C++.

170k – 240k/yr
Hybrid5+ YOEDevOps / SRE
Sigma

Senior Software Engineer - Observability and Reliability

SigmaSan Francisco, CA +1

Build observability tools and platforms (metrics, logging, tracing, alerting) using Go, OpenTelemetry, and Kubernetes. Requires 5+ years experience building high-quality software that other engineers use.

170k – 240k/yr
On-site5+ YOEDevOps / SRE