Skip to content

Latest DevOps / SRE jobs at Fal

Search
Location
5 jobs

Job results

Fal

Senior/Staff Kubernetes Infrastructure Engineer

FalUnited States

Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.

180k – 250k/yrRemote5+ YOEDevOps / SRE
Fal

Software Engineer, Site Reliability

FalUnited States

Seasoned SRE owning reliability of Kubernetes-based production infrastructure at scale for a generative AI platform. Responsibilities include operating clusters, CI/CD pipelines, SLOs, monitoring, automation with AI, and driving improvements via chaos engineering. Requires 5+ years production experience.

Salary not listedRemote5+ YOEDevOps / SRE
Fal

Software Engineer, Infrastructure

FalUnited States

Build and maintain infrastructure for a large fleet of GPU servers, including provisioning, health monitoring, diagnostics, recovery, storage optimization, and Linux tuning for AI workloads. Requires 3+ years managing large-scale bare-metal/cloud fleets, strong Python and deep Linux expertise.

Salary not listedRemote3+ YOEDevOps / SRE
Fal

Software Engineer, Infrastructure

FalSan Francisco, CA

Build and maintain infrastructure tooling for a large fleet of GPU servers, including provisioning, health monitoring, diagnostics, recovery, storage optimization, and Linux tuning to support AI workloads at scale. Requires 3+ years managing large server fleets, strong Python and deep Linux expertise.

180k – 250k/yrHybrid5+ YOEDevOps / SRE
Fal

Software Engineer, Site Reliability

FalSan Francisco, CA

Seasoned SRE owning reliability of Kubernetes-based production infrastructure at scale for a generative AI platform. Responsibilities include operating clusters, CI/CD, SLOs, monitoring, automation with AI, and driving improvements via chaos engineering. Requires 5+ years production experience with deep Kubernetes and observability expertise.

180k – 250k/yrHybrid5+ YOEDevOps / SRE