Customer Reliability Engineer
Own reliability, SLAs, and escalations for customer AI/HPC workloads at massive scale. Debug full-stack issues (hardware to scheduler), deliver technical customer incident communications, and drive root-cause fixes with internal engineering teams.
About the job
Role Scope
- Own reliability for named customer workloads: their clusters, their SLAs, their escalations.
- Debug across the full stack, hardware to fabric to scheduler, when a training run degrades.
- Run customer-facing incident communication with technical depth and no spin.
- Turn recurring customer pain into engineering fixes with the production teams.
What We're Looking For
- Supported large-scale compute customers (HPC, cloud, or AI labs) at a technical level.
- Debug distributed systems methodically across layers you don't own.
- Written incident updates customers trusted more after reading.
- Push internal teams to fix causes, not symptoms, and follow up until they do.
Bonus
- GPU training workloads.
- InfiniBand or RoCE.
- Slurm or Kubernetes.
- NCCL debugging.
Skills
Distributed Systems Debugging, Hpc, Gpu Training, InfiniBand, Roce, Slurm, Kubernetes, Nccl
Similar jobs
Support Engineering jobsOwn post-launch technical support for partner integrations by debugging production issues, improving observability, and building scalable tooling, documentation, and self-service processes. The role requires 3+ years in a client-facing technical position and strong programming, API, and troubleshooting skills.
Provides high-priority technical support to Premium and enterprise customers, troubleshooting complex platform issues, coordinating incidents, and improving support tooling and processes. Requires at least 3 years of technical support or systems engineering experience plus strong JavaScript or Python debugging skills.
Provides high-priority technical support to Premium and enterprise customers, troubleshooting complex platform issues, coordinating incidents, and improving support operations. Requires at least three years of technical support or systems engineering experience plus strong JavaScript or Python debugging skills.
Own technical customer support by troubleshooting product and account issues, writing actionable bug reports, and partnering with Engineering and Product. The role also builds scalable support systems, reporting, knowledge resources, and AI-assisted workflows for enterprise customers.
Works directly with customers to optimize API usage, deeply understanding their products and use cases while acting as an API power user to drive impact and scale.