# Network Operations Center Specialist

**Company:** [xAI](https://hotfix.jobs/companies/xai)
**Location:** Southaven, MS, Memphis, TN
**Role:** Support Engineering
**Skills:** Incident Management, Incident Response, Network Monitoring, Cluster Monitoring, Storage Monitoring, Sla Management, Runbooks, Escalation Matrices, Incident Bridges, Linear, SRE, Shift Handoffs
**Posted:** 2026-09-03

> Monitors campus infrastructure signals, triages and escalates incidents, leads incident communications, and tracks corrective actions to completion. The role requires 24/7 operations experience, strong judgment and communication skills, and rotating shift availability.

## Job Description

## Responsibilities
- Staff the NOC console on a rotating shift schedule and monitor campus signals, including cluster health, node availability, network health, facility trends, storage alarms, and threshold breaches.
- Acknowledge, classify, document, verify, and escalate pages within SLA using the escalation matrix.
- Open and run incident bridges, provide stakeholder updates on a fixed cadence, maintain live incident timelines, and identify ownership stalls.
- Prepare first-pass incident framing and hand off detailed root-cause analysis to SRE or Hardware Failure Analysis.
- Conduct structured shift handoffs and maintain durable shift logs and cross-site awareness.
- Write major-incident reports and create and track corrective projects in Linear through closure.
- Maintain and improve NOC runbooks, escalation matrices, and communications templates; participate in SRE-led game days.

## Requirements
- Experience in a 24/7 operations environment such as a NOC, SOC, dispatch, or mission control.
- Experience acknowledging, classifying, and escalating incidents under SLA.
- Experience running incident bridges, providing scheduled stakeholder updates, and maintaining incident timelines.
- Excellent written and verbal communication skills.
- Pattern recognition across compute, network, storage, and/or facilities signals.
- Experience following, maintaining, and improving operational processes such as runbooks, escalation matrices, and handoffs.
- Ability to work rotating shifts, including nights and weekends.

## Nice-to-haves
- NOC, data center operations, or campus reliability experience in high-performance computing, AI/ML infrastructure, or large-scale production environments.
- Experience writing major-incident reports and driving corrective actions to completion.
- Familiarity with Linear or similar work-tracking tools.
- Experience partnering with SRE, SiteOps, and Facilities on escalations and post-incident follow-through.
- Participation in game days, tabletop exercises, or runbook improvement programs.
- Experience at a fast-paced startup or technology company.

## Similar jobs

- [Channel Specialist](https://hotfix.jobs/jobs/09e39fb6-fb60-43dc-bd71-b0dbe5bada4e) - Fareharbor - Honolulu, HI - $43k – $64k/yr
- [Premium Support Engineer (Weekend Shift)](https://hotfix.jobs/jobs/3ea1be4f-f7cf-42e9-a614-0a571fb0e14a) - Replit - New York, NY - $185k – $210k/yr
- [Premium Support Engineer (Weekend Shift)](https://hotfix.jobs/jobs/45fc0a4b-6c07-4a47-921f-fc0ef736f887) - Replit - Foster City, CA - $185k – $210k/yr
- [Platform Support Engineer](https://hotfix.jobs/jobs/7151145d-c75b-47a2-8f77-388c9d2d75f6) - Braintrust - San Francisco, CA
- [Field Support Representative](https://hotfix.jobs/jobs/0f87aeb4-726b-48c6-9b3f-9cc2b1572175) - Skydio - Remote - $100k – $125k/yr

**Apply:** https://hotfix.jobs/jobs/bfd0e698-b582-4218-9ba5-2d3ada90b0cf
**Canonical:** https://hotfix.jobs/jobs/bfd0e698-b582-4218-9ba5-2d3ada90b0cf