Skip to content

Latest DevOps / SRE jobs at Lightning AI

Search
Location
5 jobs

Job results

Lightning AI

Senior Network Engineer

Lightning AINew York, NY +2

Senior Network Engineer responsible for designing, deploying, and optimizing large-scale NVIDIA InfiniBand fabrics and UFM for AI/ML GPU clusters. Requires 10+ years data center networking experience with deep expertise in InfiniBand, spine-leaf architectures, automation, and HPC environments.

170k – 210k/yrOn-site10+ YOEDevOps / SRE
Lightning AI

Senior Network Engineer

Lightning AIUnited States

Design and operate large-scale AI data center networks using spine-leaf architectures, EVPN/VXLAN, BGP, and automation tools. Requires 5+ years of data center networking experience and hands-on work with Cumulus NOS, SONiC, and Junos.

150k – 190k/yrRemote5+ YOEDevOps / SRE
Lightning AI

Infrastructure Operations Engineer

Lightning AINew York, NY +2

Design, build, and maintain infrastructure platforms using Linux, AWS, Kubernetes, Terraform, and Ansible to support internal and customer-facing services. Participate in on-call rotations and collaborate across engineering and operations teams.

160k – 200k/yrHybrid8+ YOEDevOps / SRE
Lightning AI

Infrastructure Engineer (GPU & Compute)

Lightning AINew York, NY +2

Owns GPU diagnostics, validation workflows, and automation for bare-metal infrastructure supporting AI/ML workloads. Requires 5+ years in systems engineering with strong Linux, Python, and NVIDIA tools expertise.

180k – 200k/yrRemote5+ YOEDevOps / SRE
Lightning AI

Infrastructure Engineer (Storage)

Lightning AINew York, NY +2

Operate and scale distributed storage systems like VAST and Ceph for AI/ML workloads, build Python automation tools, manage Linux bare-metal systems, and collaborate on infrastructure optimizations. Requires 5+ years in infrastructure engineering with storage expertise.

180k – 200k/yrRemote5+ YOEDevOps / SRE