# AI Infrastructure Systems Engineer

**Company:** [Together AI](https://hotfix.jobs/companies/together-ai)
**Location:** Bengaluru, India
**Role:** DevOps / SRE
**Experience:** 3+ years
**Skills:** Python, Go, Rust, Linux, Kubernetes, Terraform, Ansible, CUDA, Nccl, Nvlink/Nvswitch, InfiniBand, Roce, Distributed Systems, Bare-Metal Provisioning, Distributed Storage
**Posted:** 2026-08-04

> Build and operate automation, monitoring, validation, and remediation systems for a large-scale GPU fleet supporting AI training and inference. The role requires 3+ years of distributed systems or infrastructure software experience and strong Python, Go, or Rust skills.

## Job Description

## Responsibilities
- Design and build fleet automation systems to provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention.
- Build AI infrastructure agents that automate deployment, root-cause failure analysis, incident triage, and autonomous remediation.
- Develop fleet intelligence platforms to monitor hardware health, firmware, networking, storage, thermals, and workload performance, helping predict failures before they affect customers.
- Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators.
- Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads.
- Build internal platforms and developer tools for software-managed infrastructure.
- Improve deployment velocity, reliability, and operational efficiency through automation.
- Partner with hardware, networking, platform, and AI teams on large-scale AI infrastructure.

## Requirements
- 3+ years of experience building distributed systems, infrastructure platforms, or large-scale backend software.
- Strong software engineering skills in Python, Go, or Rust.
- Experience building platforms, automation systems, or developer infrastructure.
- Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies.
- Strong systems thinking across hardware and software.
- Passion for solving complex infrastructure challenges through software.
- Automation-first mindset, with an instinct to build systems that eliminate repeated manual work.

## Nice-to-Haves
- GPU infrastructure, CUDA, NCCL, or NVLink/NVSwitch.
- InfiniBand or RoCE networking.
- Bare-metal provisioning and lifecycle management.
- Large-scale AI training or inference clusters.
- Hardware health monitoring and predictive failure detection.
- Distributed storage systems.
- AI agents and autonomous infrastructure operations.

## Compensation and Benefits
- The role focuses on building infrastructure that deploys, monitors, diagnoses, optimizes, and heals GPU fleets at massive scale.

## Similar jobs

- [Software Engineer, Infrastructure](https://hotfix.jobs/jobs/585da47a-e02d-4c63-9816-248a2faa9b5b) - Granica - Remote
- [DevOps](https://hotfix.jobs/jobs/d1d4b180-df1a-4699-b783-d501b31510b4) - Acryldata - Remote
- [Site Reliability Engineer](https://hotfix.jobs/jobs/7db1f7e8-8d55-478b-9857-eb3bb0901fa0) - Invisible Tech - Remote
- [Platform Engineer - Compute Capacity](https://hotfix.jobs/jobs/4d270716-1d19-4d06-a9a3-4fdc533a14c4) - Supabase - Remote
- [Production Support Engineer](https://hotfix.jobs/jobs/328c6ec8-1e18-42d6-99c0-515d7801c9ac) - Alpaca - Remote

**Apply:** https://hotfix.jobs/jobs/0f227e06-b1bc-4a6d-99be-9dac09f43f41
**Canonical:** https://hotfix.jobs/jobs/0f227e06-b1bc-4a6d-99be-9dac09f43f41