Skip to content
xAIxAI

Network Engineer

Designs, deploys, and operates high-performance networks powering AI supercomputer campuses, including training fabrics, storage, OT, and site networks. Requires substantial data-center networking experience, automation expertise, and readiness for on-call, hands-on infrastructure work.

About the job

Responsibilities

  • Design and implement highly available, low-latency, high-bandwidth networks for AI training fabrics, inference front ends, storage, site/OT networks, and campus infrastructure.
  • Design, maintain, and operate supercomputer data center and campus networks in collaboration with infrastructure, compute, storage, SiteOps, facilities, and enterprise teams.
  • Evaluate, procure, and deploy data-center networking hardware, including switches, NICs, firewalls, optical multiplexers, and related appliances supporting 400G/800G and beyond.
  • Develop network automation tooling, including configuration analysis, linting, validation, GitOps, Infrastructure as Code, and scalable deployment frameworks.
  • Coordinate change windows for software updates, hardware refreshes, cluster expansions, and maintenance, including evenings and weekends when required.
  • Troubleshoot network issues affecting cluster health and job performance; document root-cause analyses and lead retrospectives.
  • Support cluster bring-up, expansion, and production training/inference campaigns; participate in on-call rotations.
  • Build network monitoring and telemetry for fabric health, congestion, packet loss, and NCCL/collective performance.
  • Create and maintain network architecture documentation, design drawings, fiber and cable plant records, and operational procedures.
  • Identify systemic failure modes and false redundancy across AI fabrics and site networks.
  • Gather requirements and develop implementation plans for new halls, rows, and campus interconnects with customers, vendors, and contractors.
  • Maintain network compliance with ITAR, ISO, NIST, and cybersecurity standards, including segmentation between compute, storage, OT/controls, and corporate networks.

Requirements

  • Bachelor’s degree in computer science, computer engineering, or another STEM discipline and 3+ years of professional network engineering experience, or 5+ years of professional network engineering experience in lieu of a degree.
  • Hands-on experience designing, deploying, supporting, and troubleshooting Layer 2 and Layer 3 networks in latency-sensitive, industrial, or data-center environments.
  • Experience with multiple network vendors in production or lab environments.
  • Experience using and contributing to GitOps and Infrastructure as Code frameworks.
  • Understanding of the OSI model and network standards.
  • Ability to pass applicable background checks; work in tight quarters and at heights; lift 30 pounds; drive with a valid license; and travel up to 20%.
  • Availability for extended hours, weekends, emergency 24x7 support, and after-hours on-call rotations.

Preferred Skills and Experience

  • Cisco, Arista, Juniper, or NVIDIA Spectrum-X data-center switches.
  • RoCEv2 Ethernet AI/HPC fabrics; InfiniBand experience.
  • AI training and inference traffic patterns, collectives, congestion, ECMP, adaptive routing, and NCCL.
  • WDM and large-scale single-mode or multimode fiber plants, including OTDR and acceptance testing.
  • Switch port security, network segmentation, QoS, multicast, and redundancy protocols.
  • Network monitoring, Layer 1 test tools, operational telemetry, and dashboards.
  • Bash, PowerShell, Python, Terraform, and Ansible.
  • Linux and Windows system administration.
  • CCNA or CCNP certification.
  • Real-time systems, industrial control/OT networks, or high-reliability environments in data centers, energy, aerospace, defense, or similar industries.
  • Strong communication with internal and external customers, vendors, and management.

Skills

Layer 2 Networking, Layer 3 Networking, Cisco, Arista, Juniper, Nvidia Spectrum-X, Rocev2, InfiniBand, GitOps, Infrastructure As Code, Python, Terraform, Ansible, Nccl, Linux

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Teleport

Teleport

United States

IT Security and Automation Engineer
$149k+/yrRemoteDevOps / SRE

Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.

Beacon AI

Beacon AI

San Carlos, CA

Software Engineer, Cloud Infrastructure
$135k+/yrHybridDevOps / SRE

Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.