AI Infrastructure Systems Engineer
Build and operate automation, monitoring, validation, and remediation systems for a large-scale GPU fleet supporting AI training and inference. The role requires 3+ years of distributed systems or infrastructure software experience and strong Python, Go, or Rust skills.
About the job
Responsibilities
- Design and build fleet automation systems to provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention.
- Build AI infrastructure agents that automate deployment, root-cause failure analysis, incident triage, and autonomous remediation.
- Develop fleet intelligence platforms to monitor hardware health, firmware, networking, storage, thermals, and workload performance, helping predict failures before they affect customers.
- Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators.
- Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads.
- Build internal platforms and developer tools for software-managed infrastructure.
- Improve deployment velocity, reliability, and operational efficiency through automation.
- Partner with hardware, networking, platform, and AI teams on large-scale AI infrastructure.
Requirements
- 3+ years of experience building distributed systems, infrastructure platforms, or large-scale backend software.
- Strong software engineering skills in Python, Go, or Rust.
- Experience building platforms, automation systems, or developer infrastructure.
- Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies.
- Strong systems thinking across hardware and software.
- Passion for solving complex infrastructure challenges through software.
- Automation-first mindset, with an instinct to build systems that eliminate repeated manual work.
Nice-to-Haves
- GPU infrastructure, CUDA, NCCL, or NVLink/NVSwitch.
- InfiniBand or RoCE networking.
- Bare-metal provisioning and lifecycle management.
- Large-scale AI training or inference clusters.
- Hardware health monitoring and predictive failure detection.
- Distributed storage systems.
- AI agents and autonomous infrastructure operations.
Compensation and Benefits
- The role focuses on building infrastructure that deploys, monitors, diagnoses, optimizes, and heals GPU fleets at massive scale.
Skills
Python, Go, Rust, Linux, Kubernetes, Terraform, Ansible, CUDA, Nccl, Nvlink/Nvswitch, InfiniBand, Roce, Distributed Systems, Bare-Metal Provisioning, Distributed Storage
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Own reliability, scalability, and operational excellence for DataHub Cloud and enterprise deployment offerings. The role requires 5+ years in DevOps, platform engineering, or SRE, with expertise in cloud platforms, Kubernetes, infrastructure as code, observability, and deployment automation.
Provides first-response incident triage and infrastructure stabilization for a production platform in a 24/7 rotation. Requires enterprise experience with Kubernetes, RabbitMQ, PostgreSQL, Azure, production troubleshooting, log-based diagnosis, and calm incident communication.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.