Skip to content
xAIxAI

Network Operations Center Specialist

Monitors campus infrastructure signals, triages and escalates incidents, leads incident communications, and tracks corrective actions to completion. The role requires 24/7 operations experience, strong judgment and communication skills, and rotating shift availability.

About the job

Responsibilities

  • Staff the NOC console on a rotating shift schedule and monitor campus signals, including cluster health, node availability, network health, facility trends, storage alarms, and threshold breaches.
  • Acknowledge, classify, document, verify, and escalate pages within SLA using the escalation matrix.
  • Open and run incident bridges, provide stakeholder updates on a fixed cadence, maintain live incident timelines, and identify ownership stalls.
  • Prepare first-pass incident framing and hand off detailed root-cause analysis to SRE or Hardware Failure Analysis.
  • Conduct structured shift handoffs and maintain durable shift logs and cross-site awareness.
  • Write major-incident reports and create and track corrective projects in Linear through closure.
  • Maintain and improve NOC runbooks, escalation matrices, and communications templates; participate in SRE-led game days.

Requirements

  • Experience in a 24/7 operations environment such as a NOC, SOC, dispatch, or mission control.
  • Experience acknowledging, classifying, and escalating incidents under SLA.
  • Experience running incident bridges, providing scheduled stakeholder updates, and maintaining incident timelines.
  • Excellent written and verbal communication skills.
  • Pattern recognition across compute, network, storage, and/or facilities signals.
  • Experience following, maintaining, and improving operational processes such as runbooks, escalation matrices, and handoffs.
  • Ability to work rotating shifts, including nights and weekends.

Nice-to-haves

  • NOC, data center operations, or campus reliability experience in high-performance computing, AI/ML infrastructure, or large-scale production environments.
  • Experience writing major-incident reports and driving corrective actions to completion.
  • Familiarity with Linear or similar work-tracking tools.
  • Experience partnering with SRE, SiteOps, and Facilities on escalations and post-incident follow-through.
  • Participation in game days, tabletop exercises, or runbook improvement programs.
  • Experience at a fast-paced startup or technology company.

Skills

Incident Management, Incident Response, Network Monitoring, Cluster Monitoring, Storage Monitoring, Sla Management, Runbooks, Escalation Matrices, Incident Bridges, Linear, SRE, Shift Handoffs

Fareharbor

Fareharbor

Honolulu, HI

Channel Specialist
$43k+/yrHybridSupport Engineering

Supports clients and reseller partners by managing affiliate channels, implementing partner integrations, troubleshooting invoicing and booking issues, and improving support processes. Requires strong communication, analytical problem-solving, technical comfort, and customer-service skills.

Replit

Replit

New York, NY

Premium Support Engineer (Weekend Shift)
$185k+/yrHybrid3+ YOESupport Engineering

Provides high-priority technical support to Premium and enterprise customers, troubleshooting complex platform issues, coordinating incidents, and improving support tooling and processes. Requires at least 3 years of technical support or systems engineering experience plus strong JavaScript or Python debugging skills.

Replit

Replit

Foster City, CA

Premium Support Engineer (Weekend Shift)
$185k+/yrHybrid3+ YOESupport Engineering

Provides high-priority technical support to Premium and enterprise customers, troubleshooting complex platform issues, coordinating incidents, and improving support operations. Requires at least three years of technical support or systems engineering experience plus strong JavaScript or Python debugging skills.

Braintrust

Braintrust

San Francisco, CA
Platform Support Engineer
No salary listedHybridSupport Engineering

Supports hybrid and self-hosted customer deployments by troubleshooting Kubernetes, cloud infrastructure, networking, backend performance, and reliability issues. The role combines technical customer support, incident response, coding fixes, and diagnostic tooling across major cloud platforms.

Skydio

Skydio

United States

Field Support Representative
$100k+/yrRemote3+ YOESupport Engineering

Provides advanced technical, operational, and field support for Skydio UAS hardware, docks, cloud systems, and networking. The role requires at least three years of UAS flight experience, strong troubleshooting skills, and regional travel of up to 30–50%.