Senior Network Engineer
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
About the job
Responsibilities
- Design and deploy scalable spine-leaf network architectures for AI data centers.
- Engineer high-performance Ethernet fabrics supporting GPU clusters and AI workloads.
- Build and maintain EVPN/VXLAN, BGP, and high-speed routing environments.
- Optimize east-west traffic flows for AI training and inference operations.
- Support RoCE/RDMA networking and low-latency transport technologies.
- Support backbone, data center interconnect (DCI), WAN, and edge connectivity solutions.
- Collaborate with compute, storage, AI platform, and operations teams to deliver integrated infrastructure solutions.
- Develop automation and Infrastructure-as-Code solutions for network provisioning and operations.
- Troubleshoot complex network, performance, and congestion issues across distributed environments.
- Improve network observability, telemetry, and operational visibility.
- Design systems that scale to thousands of nodes.
- Maintain strong technical documentation and communication.
Requirements
- 5+ years of experience in large-scale data center networking.
- Experience with Cumulus NOS, SONiC, and Junos.
- Experience with spine-leaf architectures and Layer 3 fabrics.
- Experience with BGP, EVPN, and VXLAN.
- Experience operating high-performance computing or GPU-dense environments.
- Experience designing networks for hyperscalers, neoclouds, or high-scale SaaS infrastructure.
- Experience with network automation using Python, Ansible, Terraform, or similar tools.
- Experience with network observability tooling and telemetry pipelines.
- Experience with cloud networking technologies, including VPCs, NFV, Direct Connect, and Cloud Connect.
- Strong documentation and communication skills.
Nice-to-Haves
- Familiarity with NVIDIA networking technologies, including Spectrum, Quantum, and BlueField.
- Familiarity with RDMA, RoCE, or InfiniBand fabrics.
- Experience with multi-region backbone design.
- Exposure to bare-metal provisioning systems.
- Experience working in high-growth infrastructure startups.
Compensation and Benefits
- Annual base salary range: $150,000–$190,000 USD.
- Discretionary bonus and meaningful equity component.
- Comprehensive medical, dental, and vision coverage.
- 401(k) matching in the United States and pension contributions in the United Kingdom.
- Unlimited paid time off, company holidays, and floating holidays.
- Two-week company-wide winter break.
- Paid parental and family leave.
- Annual learning and development allowance.
- Wellness and work-from-home stipends.
- Four weeks of paid sabbatical leave after four years of service.
- Flexible schedules and hybrid work options for office-based teams.
- Complimentary meals at office hubs.
Skills
Cumulus Nos, Sonic, Junos, BGP, Evpn, Vxlan, Roce, Rdma, Python, Ansible, Terraform, Network Telemetry, Cloud Networking, InfiniBand
Similar jobs
DevOps / SRE jobsOwn and scale infrastructure for agent orchestration, sandboxing, and hosted MCP services. The role requires hands-on Kubernetes, cloud, and infrastructure-as-code experience, along with strong software engineering fundamentals and high ownership.
Leads hybrid cloud and on-premises IT operations, incident management, automation, security hardening, and infrastructure reliability while mentoring systems engineers. Requires extensive Linux administration, ITIL operations, cloud migration, automation, and AI/ML infrastructure experience.
Senior Site Reliability Engineer providing technical leadership for scalable operations, automation, monitoring, resiliency, and cloud infrastructure. Requires a bachelor's degree, software development or architecture experience, and hands-on DevOps or systems administration experience.
Build internal developer platforms, reusable services, and automation that improve software delivery, infrastructure self-service, reliability, and developer productivity. The role requires 5+ years of platform, software, infrastructure, DevOps, or SRE experience plus expertise in cloud-native technologies, Kubernetes, CI/CD, and Infrastructure as Code.
Senior Site Reliability Engineer responsible for operating and improving large-scale, FedRAMP-compliant cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, software engineering, and reliability engineering expertise.