Latest DevOps / SRE jobs at Alembic
Job results
Design, operate, and automate the global network and reliability layer for a high-performance NVIDIA DGX SuperPOD supporting ML workloads. Own architecture, observability, incident response, and security for mission-critical infrastructure.
Builds and maintains scalable infrastructure for real-time analytics and ML workloads, focusing on reliability, automation, CI/CD, monitoring, and incident response. Requires 8+ years SRE/DevOps experience with Kubernetes, Terraform, Linux, and observability tools.