Stand up and own a 24/7 fleet operations center and end-to-end incident management for massive-scale AI data centers, including runbooks, postmortems, and driving down key metrics. Requires prior NOC/GOC leadership, structured multi-incident handling, and impactful postmortems.
188k – 237k/yr
On-site5+ YOESupport Engineering
About the role
Role Scope
Stand up and run the fleet operations center that watches every site 24/7: alarms, tickets, escalations, and communications.
Own the incident management process end to end, from first alert to postmortem, across facility and compute events.
Write the runbooks, escalation trees, and severity definitions the whole fleet operates on.
Drive incident metrics (time to acknowledge, time to resolve, repeat rate) down with process and tooling, not headcount.
What We're Looking For
You've run a NOC, GOC, or mission-control function and owned its performance numbers.
You've written incident processes that other people still use after you left.
You stay structured when several things break at once.
You write postmortems that change how the organization operates, not just what it apologizes for.
Bonus: Data center or utility operations center experience. PagerDuty or ServiceNow-class tooling. SRE-style incident frameworks.
Owns end-to-end execution of AI agent builds for enterprise customers, configuring agents, validating integrations, and collaborating with stakeholders to deliver scalable solutions. Requires 5+ years in technical customer-facing roles with strong coding and API skills.
175k – 230k/yr
On-site5+ YOESupport Engineering
Customer Reliability Engineer
FluidstackSan Francisco, CA
Own reliability, SLAs, and escalations for customer AI/HPC workloads at massive scale. Debug full-stack issues (hardware to scheduler), deliver technical customer incident communications, and drive root-cause fixes with internal engineering teams.
204k – 284k/yr
On-site5+ YOESupport Engineering
Growth Engineer, Support Engineering
FirecrawlSan Francisco, CA
Build internal tools, automations, and AI-assisted workflows on the Support Engineering team to scale developer support for Firecrawl. Requires 4+ years full-stack experience building internal tools or developer-facing systems; bonus for LLM production experience.
205k – 250k/yr
Remote4+ YOESupport Engineering
Support Engineer
AnthropicSan Francisco, CA +2
Serve as the named technical support contact for strategic enterprise accounts, owning end-to-end technical issue resolution and partnering with CS, Sales, and Applied AI teams. Requires 5+ years in escalated enterprise technical support, deep API/SaaS fluency, and experience with SSO/SAML/OAuth.
210k – 250k/yr
Hybrid5+ YOESupport Engineering
Customer Support Engineer (Inference)
Together AISan Francisco, CA
Customer Support Engineer providing technical support for AI inference and fine-tuning services on GPU clusters. Requires 5+ years customer-facing technical experience with strong AI/ML and infrastructure expertise.