Skip to content
xAIxAI

Site Reliability Engineer

Leads campus-scale site reliability for a data center environment, owning observability, incident command, postmortems, runbooks, and cross-functional reliability initiatives across infrastructure and facilities. Requires a bachelor's degree or equivalent experience and at least five years in SRE, systems engineering, or large-scale operations.

About the job

Responsibilities

  • Own monitoring architecture and signal quality, including alerting, suppression, redesign, and incorporating NOC feedback.
  • Provide technical incident leadership for SEV events, including bridge coordination, timelines, and severity management.
  • Run blameless postmortems and drive corrective actions to completion.
  • Lead cross-functional reliability projects across compute, network, storage, and facility signal boundaries.
  • Build and maintain playbooks, run game days, and maintain cross-discipline dependency maps.
  • Own runbook quality jointly with the NOC.
  • Define error budgets and availability objectives at campus and service boundaries.
  • Participate in on-call rotations and incident response for SEV-class events.

Requirements

  • Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • 5+ years of experience in site reliability, systems engineering, or large-scale production operations.
  • Large-scale incident command experience and calm technical leadership during incidents.
  • Experience designing monitoring and observability at fleet or campus scale, including alert hygiene, suppression, and signal quality.
  • Experience across at least two of compute, network, storage, power, and cooling/facilities telemetry.
  • Experience writing and operating playbooks or runbooks with a 24/7 operations or NOC partner.
  • Proficiency in Python and Bash scripting for automation and analysis.
  • General experience with at least one systems language such as C, C++, Java, Go, or Rust.
  • Strong problem-solving, data-driven reliability engineering, and cross-functional collaboration skills.

Nice to Have

  • Experience with AI/ML infrastructure or supercomputing environments.
  • Hands-on experience defining and using SLOs, SLIs, and error budgets.
  • Experience running game days, dependency mapping, and closed-loop corrective action programs.
  • Familiarity with data center hardware and plant signals, including servers, GPUs, networking, power, and cooling.
  • Experience at a fast-paced startup or technology company.

Skills

Python, Bash, C, C++, Java, Go, Rust, Observability, Incident Management, SLOs, Slis, Error Budgets, Kubernetes, Data Center Operations

Cerebras Systems

Cerebras Systems

Sunnyvale, CA

Distributed Software Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Build and operate distributed infrastructure software that automates, schedules, observes, and repairs large-scale AI compute clusters. The role requires 5+ years of infrastructure or distributed-systems experience, strong Go and Python skills, and deep Kubernetes expertise.

Mercor

Mercor

San Francisco, CA
Infrastructure Engineer
$130k+/yrOn-siteDevOps / SRE

Builds and scales highly available infrastructure using AWS, Terraform, and Docker to support rapid growth and AI workloads. Collaborates with product and research teams on architectures, CI/CD, monitoring, and performance optimization.

Cloudflare

Cloudflare

Austin, TX
Systems Engineer - Database Platform
$150k+/yrHybridDevOps / SRE

Build and operate a highly available, multi-region PostgreSQL platform, developing automation, monitoring, disaster recovery, and performance tooling. Requires experience with large-scale PostgreSQL clusters, infrastructure as code, scripting, containers, and observability.

Mercor

Mercor

San Francisco, CA

Cloud Platform Engineer
$190k+/yrOn-siteDevOps / SRE

Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.

Writer

Writer

New York, NY
Infrastructure Engineer
$140k+/yrHybrid5+ YOEDevOps / SRE

Build and operate highly available infrastructure for an enterprise AI platform, spanning cloud systems, Kubernetes, automation, observability, and reliability engineering. Requires 5+ years of production infrastructure experience, strong Python or Go skills, and daily use of AI-assisted workflows.