Senior Network Engineer
Designs, operates, and troubleshoots large-scale multi-vendor networks supporting high-performance AI compute infrastructure. Requires 8+ years of production networking experience, deep routing expertise, automation skills, and hands-on RoCE or InfiniBand experience.
About the job
Responsibilities
- Design, deploy, operate, and maintain global, multi-vendor, multi-protocol networks supporting high-performance AI compute infrastructure.
- Troubleshoot complex network and application-connectivity issues, identify root causes, and drive problems through resolution.
- Analyze telemetry, packet captures, logs, and performance data to identify network degradation, congestion, packet loss, and capacity constraints.
- Participate in architecture and design reviews for performance, availability, scalability, security, and operational supportability.
- Develop and maintain automation, validation, and operational tooling to improve network reliability and reduce manual effort.
- Evaluate network hardware, software, optics, and emerging technologies for production use.
- Establish standards and operational best practices for network design, deployment, monitoring, change management, and incident response.
- Lead projects addressing complex technical challenges and contribute to the network engineering roadmap.
- Partner with infrastructure, systems, security, and application teams on cross-functional troubleshooting.
Requirements
- 8+ years of professional experience designing, building, and supporting large-scale production data center, cloud, service-provider, or high-performance computing networks, excluding enterprise networks.
- Deep understanding of TCP/IP and experience with BGP, OSPF, VXLAN, EVPN, ECMP, and QoS.
- Experience with multi-tenant network environments using VRFs, VLANs, overlays, and policy-based segmentation.
- Hands-on experience deploying and troubleshooting network platforms from Arista, Cisco, Juniper, and NVIDIA.
- Strong troubleshooting skills using Wireshark, tcpdump, MTR, curl, nmap, and Linux networking utilities.
- Ability to diagnose connectivity, latency, packet-loss, routing, and performance issues across network, host, and application layers.
- Experience developing or maintaining network automation using Python, Ansible, or similar tools.
- Experience with Git-based software development workflows, including code review, validation, testing, CI/CD, deployment, and rollback.
- Working knowledge of Kubernetes networking, including pods, services, CNIs, and connectivity troubleshooting.
- Foundational knowledge of RDMA networking, including RoCE or InfiniBand.
- Experience with cloud networking in AWS, GCP, or Azure.
- Strong Linux administration and troubleshooting skills.
- Hands-on experience deploying or operating RoCE and/or InfiniBand fabrics.
- Experience supporting GPU clusters, HPC environments, distributed storage, or other high-bandwidth, latency-sensitive workloads.
- Understanding of AI training and inference traffic patterns.
- Experience operating networks spanning thousands of devices, multiple data centers, and geographic regions.
- Familiarity with AI-assisted engineering tools and safely validating and deploying AI-generated automation or code.
Skills
TCP/IP, BGP, Ospf, Vxlan, Evpn, Ecmp, Qos, Vrfs, Vlans, Python, Ansible, Kubernetes, Rdma, Roce, InfiniBand
Similar jobs
DevOps / SRE jobsBuild and operate multi-cloud, multi-cluster infrastructure and platform primitives for large-scale simulations and enterprise AI workloads. The role requires 5+ years in infrastructure, platform, SRE, or DevOps systems, strong Kubernetes and cloud expertise, production programming skills, and Infrastructure as Code experience.
Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.
Senior SRE who embeds with product teams to improve reliability, observability, performance, and incident preparedness. The role requires SRE or DevOps experience, strong PostgreSQL and Temporal expertise, and familiarity with observability platforms and OpenTelemetry.
Senior engineer owning safety-critical software pipelines and infrastructure, from static and dynamic analysis through CI enforcement, dashboards, and reliability tooling. Requires an advanced technical degree, 7+ years working with large codebases, and expertise in Bazel, Python, backend infrastructure, and C++.
Senior Site Reliability Engineer responsible for production troubleshooting, incident response, observability, SLOs, automation, and permanent reliability improvements. Requires strong software engineering, SQL, debugging, cloud-application troubleshooting, and cross-functional collaboration skills.