Skip to content
FluidstackFluidstackSan Francisco, CA

Customer Reliability Engineer

Own reliability, SLAs, and escalations for customer AI/HPC workloads at massive scale. Debug full-stack issues (hardware to scheduler), deliver technical customer incident communications, and drive root-cause fixes with internal engineering teams.

204k – 284k/yr
On-site5+ YOESupport Engineering

About the role

Role Scope

  • Own reliability for named customer workloads: their clusters, their SLAs, their escalations.
  • Debug across the full stack, hardware to fabric to scheduler, when a training run degrades.
  • Run customer-facing incident communication with technical depth and no spin.
  • Turn recurring customer pain into engineering fixes with the production teams.

What We're Looking For

  • Supported large-scale compute customers (HPC, cloud, or AI labs) at a technical level.
  • Debug distributed systems methodically across layers you don't own.
  • Written incident updates customers trusted more after reading.
  • Push internal teams to fix causes, not symptoms, and follow up until they do.

Bonus

  • GPU training workloads.
  • InfiniBand or RoCE.
  • Slurm or Kubernetes.
  • NCCL debugging.

Skills

distributed systems debugginghpcgpu trainingInfiniBandroceslurmKubernetesnccl
Firecrawl

Growth Engineer, Support Engineering

FirecrawlSan Francisco, CA

Build internal tools, automations, and AI-assisted workflows on the Support Engineering team to scale developer support for Firecrawl. Requires 4+ years full-stack experience building internal tools or developer-facing systems; bonus for LLM production experience.

205k – 250k/yr
Remote4+ YOESupport Engineering
Anthropic

Support Engineer

AnthropicSan Francisco, CA +2

Serve as the named technical support contact for strategic enterprise accounts, owning end-to-end technical issue resolution and partnering with CS, Sales, and Applied AI teams. Requires 5+ years in escalated enterprise technical support, deep API/SaaS fluency, and experience with SSO/SAML/OAuth.

210k – 250k/yr
Hybrid5+ YOESupport Engineering
Fluidstack

XOC & Incident Management

FluidstackAustin, TX +3

Stand up and own a 24/7 fleet operations center and end-to-end incident management for massive-scale AI data centers, including runbooks, postmortems, and driving down key metrics. Requires prior NOC/GOC leadership, structured multi-incident handling, and impactful postmortems.

188k – 237k/yr
On-site5+ YOESupport Engineering
Forward Networks

Customer Success Engineer TS/SCI

Forward NetworksMaryland

Customer Success Engineer providing post-sales technical leadership, adoption guidance, and issue resolution for Forward's network digital twin platform to Federal customers. Requires TS/SCI clearance, 5+ years customer-facing networking experience, and strong fundamentals in networking/security.

220k – 245k/yr
On-site5+ YOESupport Engineering
Decagon

Customer Engineer, Agent Builder

DecagonNew York, NY +1

Owns end-to-end execution of AI agent builds for enterprise customers, configuring agents, validating integrations, and collaborating with stakeholders to deliver scalable solutions. Requires 5+ years in technical customer-facing roles with strong coding and API skills.

175k – 230k/yr
On-site5+ YOESupport Engineering