Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.
180k – 250k/yr
Remote5+ YOEDevOps / SRE
About the role
Responsibilities
Design, automate, validate, and deliver the complete lifecycle of customer compute environments, from provisioning through upgrades, recovery, and decommissioning.
Use AI to automate and accelerate infrastructure delivery and operations.
Provision dedicated Kubernetes and Slurm clusters tailored to customer workloads.
Build and maintain Linux images and automated OS-provisioning workflows.
Operate the NVIDIA GPU stack, including drivers, GPU Operator, NVIDIA Container Toolkit, device plugins, MIG, and GPU monitoring.
Design Kubernetes and data-center networking using Cilium/Calico, MetalLB, VLAN, VXLAN, BGP, and ECMP.
Configure distributed and shared storage for high-performance workloads.
Build monitoring, alerting, diagnostics, and automated recovery for customer environments.
Develop reusable tooling, standards, documentation, and runbooks.
Collaborate with customers and internal teams to translate workload requirements into infrastructure designs.
Requirements
5+ years of experience building and operating production Linux infrastructure.
Strong production experience with Kubernetes on bare metal, including bootstrapping, upgrades, highly available control planes, etcd, containerd, CNI, CSI, ingress, load balancing, observability, security, and troubleshooting.
Experience with Linux virtualization using KVM/QEMU, libvirt, and VFIO device passthrough.
Experience operating NVIDIA GPUs on Linux and Kubernetes, including drivers, container runtimes, device plugins, GPU Operator, and GPU telemetry.
Strong networking fundamentals, including TCP/IP, L2/L3, VLANs, routing, and packet-level troubleshooting with tcpdump and Wireshark.
Practical scripting experience.
Experience with configuration-management tools such as Ansible.
Ability to diagnose complex, cross-layer infrastructure issues.
Strong communication skills and the ability to drive technical decisions across teams.
Track record of moving quickly, taking ownership, and continuously improving systems.
Nice to Have
Production Slurm experience.
High-performance networking: NVLink/NVSwitch, InfiniBand, RoCEv2, GPUDirect RDMA, NCCL, and IMEX.
Hugepages, NUMA, and CPU pinning.
SR-IOV and DPDK.
Distributed storage: Ceph, Lustre, and Weka.
KubeVirt and OpenStack.
IPsec, WireGuard, and Tailscale.
VXLAN, BGP, and ECMP.
Bare-metal management: BMC, IPMI, Redfish, PXE/iPXE, Kickstart, and cloud-init.
Network automation with NetBox, Nautobot, and Nornir.
AI training, inference, or distributed GPU workload infrastructure.
Leads strategic production engineering initiatives that improve the reliability, scalability, observability, and security of large-scale platforms. The role requires 7+ years of relevant experience, strong coding skills, and expertise in reliability practices such as SLIs, SLOs, and incident management.
Leads the operational reliability, security, observability, deployment standards, and governance of Databricks for enterprise data workloads. Requires 12+ years in platform, SRE, or cloud data infrastructure engineering plus production Databricks experience and expertise in CI/CD, secure execution, and regulated environments.
180k – 270k/yrRemote12+ YOEDevOps / SRE
Staff Engineer, AI Productivity
HightouchUnited States
Staff-level engineer building infrastructure, tooling, and documentation to make AI coding agents dramatically more productive across the codebase. Owns agentic dev environments, MCP integrations, and agent context.
180k – 400k/yrRemote7+ YOEDevOps / SRE
Staff Software Engineer, AI Developer Tools
GustoDenver, CO +3
Staff-level engineer architecting AI-native developer tools and infrastructure to accelerate engineering velocity across Gusto. Requires 8+ years experience building production AI systems with deep expertise in LLMs, RAG, and multi-agent workflows.
180k – 245k/yrHybrid8+ YOEDevOps / SRE
Staff Infrastructure Engineer
OnebriefUnited States
Staff Infrastructure Engineer building and operating secure cloud-native and edge platforms for military collaboration software. Requires 5+ years production infrastructure experience, deep Kubernetes expertise, and ability to obtain SECRET clearance.