Skip to content
FalFal

Software Engineer, Infrastructure

Build and maintain infrastructure tooling for a large fleet of GPU servers, including provisioning, health monitoring, diagnostics, recovery, storage optimization, and Linux tuning to support AI workloads at scale. Requires 3+ years managing large server fleets, strong Python and deep Linux expertise.

About the job

Key Responsibilities

  • Build and maintain Python fleet tracking system managing full lifecycle of servers (contracting, procurement, target use, pricing, availability, health, RMAs).
  • Build server management tooling that automates provisioning, health checks, GPU diagnostics, recovery, and alerting.
  • Create and maintain metrics, dashboards, and alerting for hardware health across the fleet (GPU errors, disk failures, network issues, thermals).
  • Leverage AI extensively to build tools and automate alerting and recovery.
  • Implement and enforce OS-level security: hardening baselines, SELinux/AppArmor policies, SSH key management, vulnerability scanning, and compliance automation.
  • Manage and optimize distributed and local storage systems (NVMe arrays, NFS, parallel file systems, object storage) supporting model weights, checkpoints, and ephemeral scratch.
  • Tune Linux systems for AI workloads: kernel parameters, NUMA topology, CPU pinning, hugepages, I/O schedulers, and GPU driver stack optimization (NVIDIA drivers, CUDA, container runtimes).
  • Develop automated error detection and recovery processes.
  • Work with partners to solve technical issues.

Requirements

  • 3+ years experience managing bare-metal and cloud-based server fleets at scale (100+ nodes).
  • Strong software engineering skills in Python for production tooling.
  • Deep Linux systems knowledge: boot process, kernel tuning, networking, storage, systemd, cgroups, namespaces, performance profiling.
  • Strong experience with configuration management and infrastructure-as-code (Ansible, Terraform, cloud-init).
  • Solid understanding of storage technologies: LVM, RAID, NVMe, NFS, Lustre or GPFS, Linux I/O stack tuning.
  • Familiarity with hardware diagnostics and failure modes (GPUs, NVMe, NICs, memory).
  • Experience building internal tools or dashboards for infrastructure visibility.
  • Excellent communication and ability to drive technical decisions across teams.
  • Self-starter who executes quickly, takes ownership, and constantly seeks improvement.

Nice-to-Haves

  • Familiarity with network configuration and diagnostics (VLAN, VXLAN, ECMP, BGP, tcpdump).
  • Experience with NVIDIA GPU infrastructure: driver management, health monitoring, DCGM, NVLink/NVSwitch diagnostics, RDMA, InfiniBand/RoCEv2.
  • Experience with AMD GPUs.
  • Experience with bare metal and VM provisioning (PXE/iPXE, Kickstart, libvirt, Qemu/KVM).
  • Experience with compliance frameworks (SOC 2, ISO 27001).

Compensation

$180,000-250,000 plus equity + benefits.

Skills

Python, Linux, Ansible, Terraform, Nvidia Gpus, CUDA, Nvme, Nfs, Selinux, Dcgm, InfiniBand, Lvm, Raid

Benchling

Benchling

San Francisco, CA
Software Engineer, Platform
$173k+/yrHybrid4+ YOEDevOps / SRE

Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Mercor

Mercor

San Francisco, CA

Cloud Platform Engineer
$190k+/yrOn-siteDevOps / SRE

Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.

Ramp

Ramp

New York, NY
TLM, Production Engineering
$168k+/yrHybrid3+ YOEDevOps / SRE

Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.

Roboflow

Roboflow

New York, NY
Infrastructure Engineer
$165k+/yrRemoteDevOps / SRE

Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.