Senior Network Engineer
Senior Network Engineer responsible for designing, deploying, and optimizing large-scale NVIDIA InfiniBand fabrics and UFM for AI/ML GPU clusters. Requires 10+ years data center networking experience with deep expertise in InfiniBand, spine-leaf architectures, automation, and HPC environments.
About the job
What You’ll Do
- Design, deploy, and maintain large-scale NVIDIA InfiniBand fabrics supporting AI/ML GPU clusters.
- Deploy and administer NVIDIA Unified Fabric Manager (UFM) Enterprise for monitoring, provisioning, telemetry, and fabric health.
- Configure and optimize NVIDIA Quantum and Quantum-2 InfiniBand switches.
- Troubleshoot fabric performance issues impacting NCCL, MPI, GPUDirect RDMA, and AI training jobs.
- Implement and validate fat-tree, Dragonfly+, Clos, and spine-leaf network architectures.
- Perform firmware lifecycle management for InfiniBand switches, adapters (HCAs), and UFM infrastructure.
- Optimize congestion control, adaptive routing, QoS, and traffic engineering for high-performance GPU communication.
- Work closely with AI platform, GPU infrastructure, storage, and systems engineering teams to deploy scalable AI Factory environments.
- Automate network provisioning using Python, Ansible, Git, REST APIs, and Infrastructure-as-Code methodologies.
- Monitor network health using UFM telemetry, Prometheus, Grafana, and other observability platforms.
- Support high availability, maintenance windows, incident response, root cause analysis, and capacity planning.
- Participate in architecture reviews and define networking standards for AI infrastructure.
Required Qualifications
- 7+ years of data center networking experience.
- 3+ years supporting NVIDIA InfiniBand environments.
- Hands-on experience with NVIDIA UFM Enterprise.
- Experience deploying and operating Quantum and Quantum-2 InfiniBand switches.
- Strong understanding of InfiniBand Architecture, Subnet Manager (SM), Adaptive Routing, Congestion Control, Partition Keys (PKeys), LIDs, Queue Pairs (QP), Virtual Lanes (VL), Service Levels (SL).
- Strong Linux administration experience (Ubuntu).
- Experience with automation using Python and Ansible.
- Deep understanding of Layer 2 and Layer 3 networking.
- Experience with BGP, EVPN, VXLAN, and modern spine-leaf architectures.
- Experience with packet captures and troubleshooting using tcpdump, Wireshark, and ibdiagnet tools.
- Excellent troubleshooting and communication skills.
What You Bring
- 10+ years of experience in large-scale data center networking.
- Deep expertise in spine-leaf architectures and L3 fabrics.
- Strong experience with BGP, EVPN, VXLAN.
- Experience operating high-performance computing (HPC) or GPU-dense environments.
- Experience designing networks for hyperscalers, neoclouds, or high-scale SaaS infrastructure.
- Strong automation background (Python, Ansible, Terraform, or similar).
- Experience with network observability tooling and telemetry pipelines.
- Proven ability to design systems that scale to thousands of nodes.
- Strong documentation and architectural communication skills.
Nice to Have
- Experience with Netris and Terraform.
- Experience with multi-region backbone design.
- Exposure to bare-metal provisioning systems.
- Experience working in high-growth infrastructure startups.
Benefits and Perks
- Comprehensive medical, dental and vision coverage (U.S.); Private medical and dental insurance (U.K.)
- Retirement and financial wellness support (U.S.); Pension contribution (U.K.)
- Generous paid time off, plus holidays
- Paid parental leave
- Professional development support
- Wellness and work-from-home stipends
- Flexible work environment
Skills
InfiniBand, Ufm Enterprise, Nvidia Quantum, Nvidia Quantum-2, Subnet Manager, Adaptive Routing, Congestion Control, Python, Ansible, Linux, BGP, Evpn, Vxlan, Prometheus, Grafana
Similar jobs
DevOps / SRE jobsOwn foundational cloud infrastructure and the internal developer platform supporting Commure’s engineering teams. The role requires 6+ years of infrastructure, platform, or SRE experience and hands-on expertise across Kubernetes, infrastructure as code, GitOps, observability, and cloud environments.
Leads cloud infrastructure, platform strategy, deployment pipelines, and infrastructure automation for a growing consumer platform. Requires 5+ years in infrastructure, DevOps, platform engineering, or SRE, plus deep AWS, coding, containerization, and infrastructure-as-code experience.
Own reliability, deployments, observability, compliance, and AI infrastructure across AWS and Kubernetes for a fintech platform. The role requires strong DevOps/SRE depth, backend software engineering experience, and hands-on ownership of SOC 2 and PCI-DSS controls.
Own and evolve secure, highly available AWS and Azure infrastructure, including Terraform automation, Kubernetes, CI/CD, observability, networking, and incident response. The role requires 7+ years of DevOps or related experience and strong cross-functional partnership across engineering and security.
Own and evolve VSCO’s AWS/EKS platform, including infrastructure as code, GitOps, CI/CD, observability, networking, and production reliability. The role requires 5+ years of hands-on infrastructure or SRE experience and strong Kubernetes, Terraform, and AWS expertise.