Latest DevOps / SRE jobs at Luma AI
Job results
Leads infrastructure reliability for large-scale GPU clusters, architects systems for training/inference scaling, and builds a high-performing engineering team. Requires deep Linux/distributed systems expertise, GPU production experience, and Kubernetes fluency.
Builds, maintains, and scales multi-cloud GPU infrastructure for AI training/inference, focusing on reliability, performance tuning, automation, and security in a fast-paced startup. Requires 8+ years SRE experience with deep Linux, cloud, and high-performance networking expertise.