# Site Reliability Engineer

**Company:** [xAI](https://hotfix.jobs/companies/xai)
**Location:** Memphis, TN, Southaven, MS
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** Python, Bash, C, C++, Java, Go, Rust, Observability, Incident Management, SLOs, Slis, Error Budgets, Kubernetes, Data Center Operations
**Posted:** 2026-09-02

> Leads campus-scale site reliability for a data center environment, owning observability, incident command, postmortems, runbooks, and cross-functional reliability initiatives across infrastructure and facilities. Requires a bachelor's degree or equivalent experience and at least five years in SRE, systems engineering, or large-scale operations.

## Job Description

## Responsibilities
- Own monitoring architecture and signal quality, including alerting, suppression, redesign, and incorporating NOC feedback.
- Provide technical incident leadership for SEV events, including bridge coordination, timelines, and severity management.
- Run blameless postmortems and drive corrective actions to completion.
- Lead cross-functional reliability projects across compute, network, storage, and facility signal boundaries.
- Build and maintain playbooks, run game days, and maintain cross-discipline dependency maps.
- Own runbook quality jointly with the NOC.
- Define error budgets and availability objectives at campus and service boundaries.
- Participate in on-call rotations and incident response for SEV-class events.

## Requirements
- Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience.
- 5+ years of experience in site reliability, systems engineering, or large-scale production operations.
- Large-scale incident command experience and calm technical leadership during incidents.
- Experience designing monitoring and observability at fleet or campus scale, including alert hygiene, suppression, and signal quality.
- Experience across at least two of compute, network, storage, power, and cooling/facilities telemetry.
- Experience writing and operating playbooks or runbooks with a 24/7 operations or NOC partner.
- Proficiency in Python and Bash scripting for automation and analysis.
- General experience with at least one systems language such as C, C++, Java, Go, or Rust.
- Strong problem-solving, data-driven reliability engineering, and cross-functional collaboration skills.

## Nice to Have
- Experience with AI/ML infrastructure or supercomputing environments.
- Hands-on experience defining and using SLOs, SLIs, and error budgets.
- Experience running game days, dependency mapping, and closed-loop corrective action programs.
- Familiarity with data center hardware and plant signals, including servers, GPUs, networking, power, and cooling.
- Experience at a fast-paced startup or technology company.

## Similar jobs

- [Software Engineer, Infrastructure](https://hotfix.jobs/jobs/585da47a-e02d-4c63-9816-248a2faa9b5b) - Granica - Remote
- [Capacity Ops Engineer](https://hotfix.jobs/jobs/f1904714-7dd3-4ee3-9e7a-e4fcf52083bd) - Baseten - San Francisco, CA - $170k – $230k/yr
- [IT Security and Automation Engineer](https://hotfix.jobs/jobs/604b87b5-13a2-4bba-88b2-f7d0fbbad141) - Teleport - Remote - $149k – $258k/yr
- [Electrical Field Engineer - Data Center](https://hotfix.jobs/jobs/6bfa0e4c-9ccf-438a-b65e-cd4c6297762c) - Crusoe - Remote - $196k – $235k/yr
- [Software Engineer, Cloud Infrastructure](https://hotfix.jobs/jobs/949677d6-6d57-49e8-acf8-017a14790019) - Beacon AI - San Carlos, CA - $135k – $260k/yr

**Apply:** https://hotfix.jobs/jobs/063b4870-ac81-40c6-8b50-794f4d43a2c0
**Canonical:** https://hotfix.jobs/jobs/063b4870-ac81-40c6-8b50-794f4d43a2c0