# Network Engineer

**Company:** [xAI](https://hotfix.jobs/companies/xai)
**Location:** Memphis, TN, Southaven, MS
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** Layer 2 Networking, Layer 3 Networking, Cisco, Arista, Juniper, Nvidia Spectrum-X, Rocev2, InfiniBand, GitOps, Infrastructure As Code, Python, Terraform, Ansible, Nccl, Linux
**Posted:** 2026-09-02

> Designs, deploys, and operates high-performance networks powering AI supercomputer campuses, including training fabrics, storage, OT, and site networks. Requires substantial data-center networking experience, automation expertise, and readiness for on-call, hands-on infrastructure work.

## Job Description

## Responsibilities
- Design and implement highly available, low-latency, high-bandwidth networks for AI training fabrics, inference front ends, storage, site/OT networks, and campus infrastructure.
- Design, maintain, and operate supercomputer data center and campus networks in collaboration with infrastructure, compute, storage, SiteOps, facilities, and enterprise teams.
- Evaluate, procure, and deploy data-center networking hardware, including switches, NICs, firewalls, optical multiplexers, and related appliances supporting 400G/800G and beyond.
- Develop network automation tooling, including configuration analysis, linting, validation, GitOps, Infrastructure as Code, and scalable deployment frameworks.
- Coordinate change windows for software updates, hardware refreshes, cluster expansions, and maintenance, including evenings and weekends when required.
- Troubleshoot network issues affecting cluster health and job performance; document root-cause analyses and lead retrospectives.
- Support cluster bring-up, expansion, and production training/inference campaigns; participate in on-call rotations.
- Build network monitoring and telemetry for fabric health, congestion, packet loss, and NCCL/collective performance.
- Create and maintain network architecture documentation, design drawings, fiber and cable plant records, and operational procedures.
- Identify systemic failure modes and false redundancy across AI fabrics and site networks.
- Gather requirements and develop implementation plans for new halls, rows, and campus interconnects with customers, vendors, and contractors.
- Maintain network compliance with ITAR, ISO, NIST, and cybersecurity standards, including segmentation between compute, storage, OT/controls, and corporate networks.

## Requirements
- Bachelor’s degree in computer science, computer engineering, or another STEM discipline and 3+ years of professional network engineering experience, or 5+ years of professional network engineering experience in lieu of a degree.
- Hands-on experience designing, deploying, supporting, and troubleshooting Layer 2 and Layer 3 networks in latency-sensitive, industrial, or data-center environments.
- Experience with multiple network vendors in production or lab environments.
- Experience using and contributing to GitOps and Infrastructure as Code frameworks.
- Understanding of the OSI model and network standards.
- Ability to pass applicable background checks; work in tight quarters and at heights; lift 30 pounds; drive with a valid license; and travel up to 20%.
- Availability for extended hours, weekends, emergency 24x7 support, and after-hours on-call rotations.

## Preferred Skills and Experience
- Cisco, Arista, Juniper, or NVIDIA Spectrum-X data-center switches.
- RoCEv2 Ethernet AI/HPC fabrics; InfiniBand experience.
- AI training and inference traffic patterns, collectives, congestion, ECMP, adaptive routing, and NCCL.
- WDM and large-scale single-mode or multimode fiber plants, including OTDR and acceptance testing.
- Switch port security, network segmentation, QoS, multicast, and redundancy protocols.
- Network monitoring, Layer 1 test tools, operational telemetry, and dashboards.
- Bash, PowerShell, Python, Terraform, and Ansible.
- Linux and Windows system administration.
- CCNA or CCNP certification.
- Real-time systems, industrial control/OT networks, or high-reliability environments in data centers, energy, aerospace, defense, or similar industries.
- Strong communication with internal and external customers, vendors, and management.

## Similar jobs

- [Software Engineer, Infrastructure](https://hotfix.jobs/jobs/585da47a-e02d-4c63-9816-248a2faa9b5b) - Granica - Remote
- [Capacity Ops Engineer](https://hotfix.jobs/jobs/f1904714-7dd3-4ee3-9e7a-e4fcf52083bd) - Baseten - San Francisco, CA - $170k – $230k/yr
- [IT Security and Automation Engineer](https://hotfix.jobs/jobs/604b87b5-13a2-4bba-88b2-f7d0fbbad141) - Teleport - Remote - $149k – $258k/yr
- [Electrical Field Engineer - Data Center](https://hotfix.jobs/jobs/6bfa0e4c-9ccf-438a-b65e-cd4c6297762c) - Crusoe - Remote - $196k – $235k/yr
- [Software Engineer, Cloud Infrastructure](https://hotfix.jobs/jobs/949677d6-6d57-49e8-acf8-017a14790019) - Beacon AI - San Carlos, CA - $135k – $260k/yr

**Apply:** https://hotfix.jobs/jobs/f9ed31f3-6e40-4ce5-9efe-474656825674
**Canonical:** https://hotfix.jobs/jobs/f9ed31f3-6e40-4ce5-9efe-474656825674