AI Infrastructure Systems Engineer
Build and operate autonomous infrastructure systems for large-scale GPU fleets, including cluster lifecycle automation, fleet intelligence, validation, and remediation. The role requires 3+ years of distributed systems or infrastructure engineering experience and strong Python, Go, or Rust skills.
About the job
Responsibilities
- Design and build fleet automation systems to provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention.
- Build AI infrastructure agents for deployment automation, root-cause analysis, incident triage, and autonomous remediation.
- Develop fleet intelligence platforms monitoring hardware health, firmware, networking, storage, thermals, and workload performance to predict failures.
- Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators.
- Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads.
- Build internal platforms and developer tools for software-managed infrastructure.
- Improve deployment velocity, reliability, and operational efficiency through automation.
- Partner with hardware, networking, platform, and AI teams.
Requirements
- 3+ years building distributed systems, infrastructure platforms, or large-scale backend software.
- Strong software engineering skills in Python, Go, or Rust.
- Experience building platforms, automation systems, or developer infrastructure.
- Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies.
- Strong systems thinking across hardware and software.
- Automation-first approach to eliminating repetitive operational work.
Nice-to-haves
- GPU infrastructure, CUDA, NCCL, and NVLink/NVSwitch.
- InfiniBand or RoCE networking.
- Bare-metal provisioning and lifecycle management.
- Large-scale AI training or inference clusters.
- Hardware health monitoring and predictive failure detection.
- Distributed storage systems.
- AI agents and autonomous infrastructure operations.
Skills
Python, Go, Rust, Linux, Kubernetes, Terraform, Ansible, CUDA, Nccl, Nvlink/Nvswitch, InfiniBand, Roce, Distributed Systems, Gpu Infrastructure, Distributed Storage
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.