Skip to content

AI Infrastructure Systems Engineer

Build and operate autonomous infrastructure systems for large-scale GPU fleets, including cluster lifecycle automation, fleet intelligence, validation, and remediation. The role requires 3+ years of distributed systems or infrastructure engineering experience and strong Python, Go, or Rust skills.

About the job

Responsibilities

  • Design and build fleet automation systems to provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention.
  • Build AI infrastructure agents for deployment automation, root-cause analysis, incident triage, and autonomous remediation.
  • Develop fleet intelligence platforms monitoring hardware health, firmware, networking, storage, thermals, and workload performance to predict failures.
  • Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators.
  • Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads.
  • Build internal platforms and developer tools for software-managed infrastructure.
  • Improve deployment velocity, reliability, and operational efficiency through automation.
  • Partner with hardware, networking, platform, and AI teams.

Requirements

  • 3+ years building distributed systems, infrastructure platforms, or large-scale backend software.
  • Strong software engineering skills in Python, Go, or Rust.
  • Experience building platforms, automation systems, or developer infrastructure.
  • Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies.
  • Strong systems thinking across hardware and software.
  • Automation-first approach to eliminating repetitive operational work.

Nice-to-haves

  • GPU infrastructure, CUDA, NCCL, and NVLink/NVSwitch.
  • InfiniBand or RoCE networking.
  • Bare-metal provisioning and lifecycle management.
  • Large-scale AI training or inference clusters.
  • Hardware health monitoring and predictive failure detection.
  • Distributed storage systems.
  • AI agents and autonomous infrastructure operations.

Skills

Python, Go, Rust, Linux, Kubernetes, Terraform, Ansible, CUDA, Nccl, Nvlink/Nvswitch, InfiniBand, Roce, Distributed Systems, Gpu Infrastructure, Distributed Storage

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Teleport

Teleport

United States

IT Security and Automation Engineer
$149k+/yrRemoteDevOps / SRE

Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.

Beacon AI

Beacon AI

San Carlos, CA

Software Engineer, Cloud Infrastructure
$135k+/yrHybridDevOps / SRE

Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.