# Senior Network Engineer

**Company:** [Lightning AI](https://hotfix.jobs/companies/lightning-ai)
**Location:** New York, NY, San Francisco, CA, Seattle, WA
**Role:** DevOps / SRE
**Salary:** $170k – $210k/yr
**Experience:** 10+ years
**Skills:** InfiniBand, ufm enterprise, nvidia quantum, nvidia quantum-2, subnet manager, adaptive routing, congestion control, Python, Ansible, Linux, BGP, evpn, vxlan, Prometheus, Grafana
**Posted:** 2026-07-20

> Senior Network Engineer responsible for designing, deploying, and optimizing large-scale NVIDIA InfiniBand fabrics and UFM for AI/ML GPU clusters. Requires 10+ years data center networking experience with deep expertise in InfiniBand, spine-leaf architectures, automation, and HPC environments.

## Job Description

## What You’ll Do
- Design, deploy, and maintain large-scale NVIDIA InfiniBand fabrics supporting AI/ML GPU clusters.
- Deploy and administer NVIDIA Unified Fabric Manager (UFM) Enterprise for monitoring, provisioning, telemetry, and fabric health.
- Configure and optimize NVIDIA Quantum and Quantum-2 InfiniBand switches.
- Troubleshoot fabric performance issues impacting NCCL, MPI, GPUDirect RDMA, and AI training jobs.
- Implement and validate fat-tree, Dragonfly+, Clos, and spine-leaf network architectures.
- Perform firmware lifecycle management for InfiniBand switches, adapters (HCAs), and UFM infrastructure.
- Optimize congestion control, adaptive routing, QoS, and traffic engineering for high-performance GPU communication.
- Work closely with AI platform, GPU infrastructure, storage, and systems engineering teams to deploy scalable AI Factory environments.
- Automate network provisioning using Python, Ansible, Git, REST APIs, and Infrastructure-as-Code methodologies.
- Monitor network health using UFM telemetry, Prometheus, Grafana, and other observability platforms.
- Support high availability, maintenance windows, incident response, root cause analysis, and capacity planning.
- Participate in architecture reviews and define networking standards for AI infrastructure.

## Required Qualifications
- 7+ years of data center networking experience.
- 3+ years supporting NVIDIA InfiniBand environments.
- Hands-on experience with NVIDIA UFM Enterprise.
- Experience deploying and operating Quantum and Quantum-2 InfiniBand switches.
- Strong understanding of InfiniBand Architecture, Subnet Manager (SM), Adaptive Routing, Congestion Control, Partition Keys (PKeys), LIDs, Queue Pairs (QP), Virtual Lanes (VL), Service Levels (SL).
- Strong Linux administration experience (Ubuntu).
- Experience with automation using Python and Ansible.
- Deep understanding of Layer 2 and Layer 3 networking.
- Experience with BGP, EVPN, VXLAN, and modern spine-leaf architectures.
- Experience with packet captures and troubleshooting using tcpdump, Wireshark, and ibdiagnet tools.
- Excellent troubleshooting and communication skills.

## What You Bring
- 10+ years of experience in large-scale data center networking.
- Deep expertise in spine-leaf architectures and L3 fabrics.
- Strong experience with BGP, EVPN, VXLAN.
- Experience operating high-performance computing (HPC) or GPU-dense environments.
- Experience designing networks for hyperscalers, neoclouds, or high-scale SaaS infrastructure.
- Strong automation background (Python, Ansible, Terraform, or similar).
- Experience with network observability tooling and telemetry pipelines.
- Proven ability to design systems that scale to thousands of nodes.
- Strong documentation and architectural communication skills.

## Nice to Have
- Experience with Netris and Terraform.
- Experience with multi-region backbone design.
- Exposure to bare-metal provisioning systems.
- Experience working in high-growth infrastructure startups.

## Benefits and Perks
- Comprehensive medical, dental and vision coverage (U.S.); Private medical and dental insurance (U.K.)
- Retirement and financial wellness support (U.S.); Pension contribution (U.K.)
- Generous paid time off, plus holidays
- Paid parental leave
- Professional development support
- Wellness and work-from-home stipends
- Flexible work environment

## Similar roles

- [Senior Performance Engineer](https://hotfix.jobs/jobs/6629d9ee-877d-4387-8d8f-e75835fd70c0) - Crusoe - San Francisco, CA - $170k – $205k/yr
- [Sr. Site Reliability Engineer](https://hotfix.jobs/jobs/ee28b13f-c7af-4be6-a7be-07ce78ec3108) - Illumio - Sunnyvale, CA - $170k – $196k/yr
- [Software Engineer, Compute Infrastructure](https://hotfix.jobs/jobs/55f94e84-3561-47ca-abbf-933b6bad3ee2) - Render - Remote - $170k – $290k/yr
- [Senior Software Engineer, Developer Productivity Cloud Infrastructure](https://hotfix.jobs/jobs/f213ddfe-a997-4ee7-af19-46182e0c937a) - Skydio - San Mateo, CA - $170k – $240k/yr
- [Senior Software Engineer - Observability and Reliability](https://hotfix.jobs/jobs/1c60bbb9-ebb3-437e-bd25-590308c243d4) - Sigma - San Francisco, CA - $170k – $240k/yr

**Apply:** https://hotfix.jobs/jobs/0e8d10d2-42aa-417b-a1a9-12e83e500931
**Canonical:** https://hotfix.jobs/jobs/0e8d10d2-42aa-417b-a1a9-12e83e500931