Skip to content

AI Infrastructure Systems Engineer

Build and operate automation, monitoring, validation, and remediation systems for a large-scale GPU fleet supporting AI training and inference. The role requires 3+ years of distributed systems or infrastructure software experience and strong Python, Go, or Rust skills.

About the job

Responsibilities

  • Design and build fleet automation systems to provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention.
  • Build AI infrastructure agents that automate deployment, root-cause failure analysis, incident triage, and autonomous remediation.
  • Develop fleet intelligence platforms to monitor hardware health, firmware, networking, storage, thermals, and workload performance, helping predict failures before they affect customers.
  • Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators.
  • Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads.
  • Build internal platforms and developer tools for software-managed infrastructure.
  • Improve deployment velocity, reliability, and operational efficiency through automation.
  • Partner with hardware, networking, platform, and AI teams on large-scale AI infrastructure.

Requirements

  • 3+ years of experience building distributed systems, infrastructure platforms, or large-scale backend software.
  • Strong software engineering skills in Python, Go, or Rust.
  • Experience building platforms, automation systems, or developer infrastructure.
  • Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies.
  • Strong systems thinking across hardware and software.
  • Passion for solving complex infrastructure challenges through software.
  • Automation-first mindset, with an instinct to build systems that eliminate repeated manual work.

Nice-to-Haves

  • GPU infrastructure, CUDA, NCCL, or NVLink/NVSwitch.
  • InfiniBand or RoCE networking.
  • Bare-metal provisioning and lifecycle management.
  • Large-scale AI training or inference clusters.
  • Hardware health monitoring and predictive failure detection.
  • Distributed storage systems.
  • AI agents and autonomous infrastructure operations.

Compensation and Benefits

  • The role focuses on building infrastructure that deploys, monitors, diagnoses, optimizes, and heals GPU fleets at massive scale.

Skills

Python, Go, Rust, Linux, Kubernetes, Terraform, Ansible, CUDA, Nccl, Nvlink/Nvswitch, InfiniBand, Roce, Distributed Systems, Bare-Metal Provisioning, Distributed Storage

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Acryldata

Acryldata

Bengaluru, India

DevOps
No salary listedRemote5+ YOEDevOps / SRE

Own reliability, scalability, and operational excellence for DataHub Cloud and enterprise deployment offerings. The role requires 5+ years in DevOps, platform engineering, or SRE, with expertise in cloud platforms, Kubernetes, infrastructure as code, observability, and deployment automation.

Invisible Tech

Invisible Tech

Estonia
Site Reliability Engineer
No salary listedRemoteDevOps / SRE

Provides first-response incident triage and infrastructure stabilization for a production platform in a 24/7 rotation. Requires enterprise experience with Kubernetes, RabbitMQ, PostgreSQL, Azure, production troubleshooting, log-based diagnosis, and calm incident communication.

Supabase

Supabase

Remote

Platform Engineer - Compute Capacity
No salary listedRemote5+ YOEDevOps / SRE

Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.

Alpaca

Alpaca

Remote

Production Support Engineer
No salary listedRemote4+ YOEDevOps / SRE

Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.