Skip to content
FalFalUnited States

Senior/Staff Kubernetes Infrastructure Engineer

Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.

180k – 250k/yr
Remote5+ YOEDevOps / SRE

About the role

Responsibilities

  • Design, automate, validate, and deliver the complete lifecycle of customer compute environments, from provisioning through upgrades, recovery, and decommissioning.
  • Use AI to automate and accelerate infrastructure delivery and operations.
  • Provision dedicated Kubernetes and Slurm clusters tailored to customer workloads.
  • Build and maintain Linux images and automated OS-provisioning workflows.
  • Operate the NVIDIA GPU stack, including drivers, GPU Operator, NVIDIA Container Toolkit, device plugins, MIG, and GPU monitoring.
  • Design Kubernetes and data-center networking using Cilium/Calico, MetalLB, VLAN, VXLAN, BGP, and ECMP.
  • Configure distributed and shared storage for high-performance workloads.
  • Build monitoring, alerting, diagnostics, and automated recovery for customer environments.
  • Develop reusable tooling, standards, documentation, and runbooks.
  • Collaborate with customers and internal teams to translate workload requirements into infrastructure designs.

Requirements

  • 5+ years of experience building and operating production Linux infrastructure.
  • Strong production experience with Kubernetes on bare metal, including bootstrapping, upgrades, highly available control planes, etcd, containerd, CNI, CSI, ingress, load balancing, observability, security, and troubleshooting.
  • Experience with Linux virtualization using KVM/QEMU, libvirt, and VFIO device passthrough.
  • Experience operating NVIDIA GPUs on Linux and Kubernetes, including drivers, container runtimes, device plugins, GPU Operator, and GPU telemetry.
  • Strong networking fundamentals, including TCP/IP, L2/L3, VLANs, routing, and packet-level troubleshooting with tcpdump and Wireshark.
  • Practical scripting experience.
  • Experience with configuration-management tools such as Ansible.
  • Ability to diagnose complex, cross-layer infrastructure issues.
  • Strong communication skills and the ability to drive technical decisions across teams.
  • Track record of moving quickly, taking ownership, and continuously improving systems.

Nice to Have

  • Production Slurm experience.
  • High-performance networking: NVLink/NVSwitch, InfiniBand, RoCEv2, GPUDirect RDMA, NCCL, and IMEX.
  • Hugepages, NUMA, and CPU pinning.
  • SR-IOV and DPDK.
  • Distributed storage: Ceph, Lustre, and Weka.
  • KubeVirt and OpenStack.
  • IPsec, WireGuard, and Tailscale.
  • VXLAN, BGP, and ECMP.
  • Bare-metal management: BMC, IPMI, Redfish, PXE/iPXE, Kickstart, and cloud-init.
  • Network automation with NetBox, Nautobot, and Nornir.
  • AI training, inference, or distributed GPU workload infrastructure.
  • Python or Go proficiency.

Skills

KubernetesLinuxslurmnvidia gpusgpu operatorkvm/qemulibvirtvfiociliumcalicoAnsiblePythonGocephBGP

Similar roles

DevOps / SRE jobs
Attentive

Staff Site Reliability Engineer

AttentiveUnited States

Leads strategic production engineering initiatives that improve the reliability, scalability, observability, and security of large-scale platforms. The role requires 7+ years of relevant experience, strong coding skills, and expertise in reliability practices such as SLIs, SLOs, and incident management.

180k – 240k/yrRemote7+ YOEDevOps / SRE
Shield AI

Sr. Staff Platform/Data Reliability Engineer, Databricks

Shield AIUnited States

Leads the operational reliability, security, observability, deployment standards, and governance of Databricks for enterprise data workloads. Requires 12+ years in platform, SRE, or cloud data infrastructure engineering plus production Databricks experience and expertise in CI/CD, secure execution, and regulated environments.

180k – 270k/yrRemote12+ YOEDevOps / SRE
Hightouch

Staff Engineer, AI Productivity

HightouchUnited States

Staff-level engineer building infrastructure, tooling, and documentation to make AI coding agents dramatically more productive across the codebase. Owns agentic dev environments, MCP integrations, and agent context.

180k – 400k/yrRemote7+ YOEDevOps / SRE
Gusto

Staff Software Engineer, AI Developer Tools

GustoDenver, CO +3

Staff-level engineer architecting AI-native developer tools and infrastructure to accelerate engineering velocity across Gusto. Requires 8+ years experience building production AI systems with deep expertise in LLMs, RAG, and multi-agent workflows.

180k – 245k/yrHybrid8+ YOEDevOps / SRE
Onebrief

Staff Infrastructure Engineer

OnebriefUnited States

Staff Infrastructure Engineer building and operating secure cloud-native and edge platforms for military collaboration software. Requires 5+ years production infrastructure experience, deep Kubernetes expertise, and ability to obtain SECRET clearance.

180k – 235k/yrRemote5+ YOEDevOps / SRE