Site Reliability Engineer
Leads campus-scale site reliability for a data center environment, owning observability, incident command, postmortems, runbooks, and cross-functional reliability initiatives across infrastructure and facilities. Requires a bachelor's degree or equivalent experience and at least five years in SRE, systems engineering, or large-scale operations.
About the job
Responsibilities
- Own monitoring architecture and signal quality, including alerting, suppression, redesign, and incorporating NOC feedback.
- Provide technical incident leadership for SEV events, including bridge coordination, timelines, and severity management.
- Run blameless postmortems and drive corrective actions to completion.
- Lead cross-functional reliability projects across compute, network, storage, and facility signal boundaries.
- Build and maintain playbooks, run game days, and maintain cross-discipline dependency maps.
- Own runbook quality jointly with the NOC.
- Define error budgets and availability objectives at campus and service boundaries.
- Participate in on-call rotations and incident response for SEV-class events.
Requirements
- Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience.
- 5+ years of experience in site reliability, systems engineering, or large-scale production operations.
- Large-scale incident command experience and calm technical leadership during incidents.
- Experience designing monitoring and observability at fleet or campus scale, including alert hygiene, suppression, and signal quality.
- Experience across at least two of compute, network, storage, power, and cooling/facilities telemetry.
- Experience writing and operating playbooks or runbooks with a 24/7 operations or NOC partner.
- Proficiency in Python and Bash scripting for automation and analysis.
- General experience with at least one systems language such as C, C++, Java, Go, or Rust.
- Strong problem-solving, data-driven reliability engineering, and cross-functional collaboration skills.
Nice to Have
- Experience with AI/ML infrastructure or supercomputing environments.
- Hands-on experience defining and using SLOs, SLIs, and error budgets.
- Experience running game days, dependency mapping, and closed-loop corrective action programs.
- Familiarity with data center hardware and plant signals, including servers, GPUs, networking, power, and cooling.
- Experience at a fast-paced startup or technology company.
Skills
Python, Bash, C, C++, Java, Go, Rust, Observability, Incident Management, SLOs, Slis, Error Budgets, Kubernetes, Data Center Operations
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.