# AI Infrastructure Systems Engineer

**Company:** [Together AI](https://hotfix.jobs/companies/together-ai)
**Location:** San Francisco, CA, Amsterdam, Netherlands
**Role:** DevOps / SRE
**Experience:** 3+ years
**Skills:** Python, Go, Rust, Linux, Kubernetes, Terraform, Ansible, CUDA, Nccl, Nvlink/Nvswitch, InfiniBand, Roce, Distributed Systems, Gpu Infrastructure, Distributed Storage
**Posted:** 2026-09-03

> Build and operate autonomous infrastructure systems for large-scale GPU fleets, including cluster lifecycle automation, fleet intelligence, validation, and remediation. The role requires 3+ years of distributed systems or infrastructure engineering experience and strong Python, Go, or Rust skills.

## Job Description

## Responsibilities
- Design and build fleet automation systems to provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention.
- Build AI infrastructure agents for deployment automation, root-cause analysis, incident triage, and autonomous remediation.
- Develop fleet intelligence platforms monitoring hardware health, firmware, networking, storage, thermals, and workload performance to predict failures.
- Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators.
- Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads.
- Build internal platforms and developer tools for software-managed infrastructure.
- Improve deployment velocity, reliability, and operational efficiency through automation.
- Partner with hardware, networking, platform, and AI teams.

## Requirements
- 3+ years building distributed systems, infrastructure platforms, or large-scale backend software.
- Strong software engineering skills in Python, Go, or Rust.
- Experience building platforms, automation systems, or developer infrastructure.
- Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies.
- Strong systems thinking across hardware and software.
- Automation-first approach to eliminating repetitive operational work.

## Nice-to-haves
- GPU infrastructure, CUDA, NCCL, and NVLink/NVSwitch.
- InfiniBand or RoCE networking.
- Bare-metal provisioning and lifecycle management.
- Large-scale AI training or inference clusters.
- Hardware health monitoring and predictive failure detection.
- Distributed storage systems.
- AI agents and autonomous infrastructure operations.

## Similar jobs

- [Software Engineer, Infrastructure](https://hotfix.jobs/jobs/585da47a-e02d-4c63-9816-248a2faa9b5b) - Granica - Remote
- [Capacity Ops Engineer](https://hotfix.jobs/jobs/f1904714-7dd3-4ee3-9e7a-e4fcf52083bd) - Baseten - San Francisco, CA - $170k – $230k/yr
- [IT Security and Automation Engineer](https://hotfix.jobs/jobs/604b87b5-13a2-4bba-88b2-f7d0fbbad141) - Teleport - Remote - $149k – $258k/yr
- [Electrical Field Engineer - Data Center](https://hotfix.jobs/jobs/6bfa0e4c-9ccf-438a-b65e-cd4c6297762c) - Crusoe - Remote - $196k – $235k/yr
- [Software Engineer, Cloud Infrastructure](https://hotfix.jobs/jobs/949677d6-6d57-49e8-acf8-017a14790019) - Beacon AI - San Carlos, CA - $135k – $260k/yr

**Apply:** https://hotfix.jobs/jobs/34b2c44e-c259-48df-a98e-0e98e0da2aab
**Canonical:** https://hotfix.jobs/jobs/34b2c44e-c259-48df-a98e-0e98e0da2aab