Skip to content
FluidstackFluidstackAustin, TX

XOC & Incident Management

Stand up and own a 24/7 fleet operations center and end-to-end incident management for massive-scale AI data centers, including runbooks, postmortems, and driving down key metrics. Requires prior NOC/GOC leadership, structured multi-incident handling, and impactful postmortems.

188k – 237k/yr
On-site5+ YOESupport Engineering

About the role

Role Scope

  • Stand up and run the fleet operations center that watches every site 24/7: alarms, tickets, escalations, and communications.
  • Own the incident management process end to end, from first alert to postmortem, across facility and compute events.
  • Write the runbooks, escalation trees, and severity definitions the whole fleet operates on.
  • Drive incident metrics (time to acknowledge, time to resolve, repeat rate) down with process and tooling, not headcount.

What We're Looking For

  • You've run a NOC, GOC, or mission-control function and owned its performance numbers.
  • You've written incident processes that other people still use after you left.
  • You stay structured when several things break at once.
  • You write postmortems that change how the organization operates, not just what it apologizes for.
  • Bonus: Data center or utility operations center experience. PagerDuty or ServiceNow-class tooling. SRE-style incident frameworks.

Skills

Incident Managementnocgocrunbooksescalation processespostmortemspagerdutyservicenowsre frameworks
Decagon

Customer Engineer, Agent Builder

DecagonNew York, NY +1

Owns end-to-end execution of AI agent builds for enterprise customers, configuring agents, validating integrations, and collaborating with stakeholders to deliver scalable solutions. Requires 5+ years in technical customer-facing roles with strong coding and API skills.

175k – 230k/yr
On-site5+ YOESupport Engineering
Fluidstack

Customer Reliability Engineer

FluidstackSan Francisco, CA

Own reliability, SLAs, and escalations for customer AI/HPC workloads at massive scale. Debug full-stack issues (hardware to scheduler), deliver technical customer incident communications, and drive root-cause fixes with internal engineering teams.

204k – 284k/yr
On-site5+ YOESupport Engineering
Firecrawl

Growth Engineer, Support Engineering

FirecrawlSan Francisco, CA

Build internal tools, automations, and AI-assisted workflows on the Support Engineering team to scale developer support for Firecrawl. Requires 4+ years full-stack experience building internal tools or developer-facing systems; bonus for LLM production experience.

205k – 250k/yr
Remote4+ YOESupport Engineering
Anthropic

Support Engineer

AnthropicSan Francisco, CA +2

Serve as the named technical support contact for strategic enterprise accounts, owning end-to-end technical issue resolution and partnering with CS, Sales, and Applied AI teams. Requires 5+ years in escalated enterprise technical support, deep API/SaaS fluency, and experience with SSO/SAML/OAuth.

210k – 250k/yr
Hybrid5+ YOESupport Engineering
Together AI

Customer Support Engineer (Inference)

Together AISan Francisco, CA

Customer Support Engineer providing technical support for AI inference and fine-tuning services on GPU clusters. Requires 5+ years customer-facing technical experience with strong AI/ML and infrastructure expertise.

160k – 230k/yr
Remote5+ YOESupport Engineering