Own reliability, SLAs, and escalations for customer AI/HPC workloads at massive scale. Debug full-stack issues (hardware to scheduler), deliver technical customer incident communications, and drive root-cause fixes with internal engineering teams.
204k – 284k/yr
On-site5+ YOESupport Engineering
About the role
Role Scope
Own reliability for named customer workloads: their clusters, their SLAs, their escalations.
Debug across the full stack, hardware to fabric to scheduler, when a training run degrades.
Run customer-facing incident communication with technical depth and no spin.
Turn recurring customer pain into engineering fixes with the production teams.
What We're Looking For
Supported large-scale compute customers (HPC, cloud, or AI labs) at a technical level.
Debug distributed systems methodically across layers you don't own.
Written incident updates customers trusted more after reading.
Push internal teams to fix causes, not symptoms, and follow up until they do.
Bonus
GPU training workloads.
InfiniBand or RoCE.
Slurm or Kubernetes.
NCCL debugging.
Skills
distributed systems debugginghpcgpu trainingInfiniBandroceslurmKubernetesnccl
Build internal tools, automations, and AI-assisted workflows on the Support Engineering team to scale developer support for Firecrawl. Requires 4+ years full-stack experience building internal tools or developer-facing systems; bonus for LLM production experience.
205k – 250k/yr
Remote4+ YOESupport Engineering
Support Engineer
AnthropicSan Francisco, CA +2
Serve as the named technical support contact for strategic enterprise accounts, owning end-to-end technical issue resolution and partnering with CS, Sales, and Applied AI teams. Requires 5+ years in escalated enterprise technical support, deep API/SaaS fluency, and experience with SSO/SAML/OAuth.
210k – 250k/yr
Hybrid5+ YOESupport Engineering
XOC & Incident Management
FluidstackAustin, TX +3
Stand up and own a 24/7 fleet operations center and end-to-end incident management for massive-scale AI data centers, including runbooks, postmortems, and driving down key metrics. Requires prior NOC/GOC leadership, structured multi-incident handling, and impactful postmortems.
188k – 237k/yr
On-site5+ YOESupport Engineering
Customer Success Engineer TS/SCI
Forward NetworksMaryland
Customer Success Engineer providing post-sales technical leadership, adoption guidance, and issue resolution for Forward's network digital twin platform to Federal customers. Requires TS/SCI clearance, 5+ years customer-facing networking experience, and strong fundamentals in networking/security.
220k – 245k/yr
On-site5+ YOESupport Engineering
Customer Engineer, Agent Builder
DecagonNew York, NY +1
Owns end-to-end execution of AI agent builds for enterprise customers, configuring agents, validating integrations, and collaborating with stakeholders to deliver scalable solutions. Requires 5+ years in technical customer-facing roles with strong coding and API skills.