Site Reliability Engineer
Leads campus-scale site reliability for a data center environment, owning observability, incident command, postmortems, runbooks, and cross-functional reliability initiatives across infrastructure and facilities. Requires a bachelor's degree or equivalent experience and at least five years in SRE, systems engineering, or large-scale operations.
About the job
Responsibilities
- Own monitoring architecture and signal quality, including alerting, suppression, redesign, and incorporating NOC feedback.
- Provide technical incident leadership for SEV events, including bridge coordination, timelines, and severity management.
- Run blameless postmortems and drive corrective actions to completion.
- Lead cross-functional reliability projects across compute, network, storage, and facility signal boundaries.
- Build and maintain playbooks, run game days, and maintain cross-discipline dependency maps.
- Own runbook quality jointly with the NOC.
- Define error budgets and availability objectives at campus and service boundaries.
- Participate in on-call rotations and incident response for SEV-class events.
Requirements
- Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent experience.
- 5+ years of experience in site reliability, systems engineering, or large-scale production operations.
- Large-scale incident command experience and calm technical leadership during incidents.
- Experience designing monitoring and observability at fleet or campus scale, including alert hygiene, suppression, and signal quality.
- Experience across at least two of compute, network, storage, power, and cooling/facilities telemetry.
- Experience writing and operating playbooks or runbooks with a 24/7 operations or NOC partner.
- Proficiency in Python and Bash scripting for automation and analysis.
- General experience with at least one systems language such as C, C++, Java, Go, or Rust.
- Strong problem-solving, data-driven reliability engineering, and cross-functional collaboration skills.
Nice to Have
- Experience with AI/ML infrastructure or supercomputing environments.
- Hands-on experience defining and using SLOs, SLIs, and error budgets.
- Experience running game days, dependency mapping, and closed-loop corrective action programs.
- Familiarity with data center hardware and plant signals, including servers, GPUs, networking, power, and cooling.
- Experience at a fast-paced startup or technology company.
Skills
Python, Bash, C, C++, Java, Go, Rust, Observability, Incident Management, SLOs, Slis, Error Budgets, Kubernetes, Data Center Operations
Similar jobs
DevOps / SRE jobsBuild and operate distributed infrastructure software that automates, schedules, observes, and repairs large-scale AI compute clusters. The role requires 5+ years of infrastructure or distributed-systems experience, strong Go and Python skills, and deep Kubernetes expertise.
Builds and scales highly available infrastructure using AWS, Terraform, and Docker to support rapid growth and AI workloads. Collaborates with product and research teams on architectures, CI/CD, monitoring, and performance optimization.
Build and operate a highly available, multi-region PostgreSQL platform, developing automation, monitoring, disaster recovery, and performance tooling. Requires experience with large-scale PostgreSQL clusters, infrastructure as code, scripting, containers, and observability.
Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.
Build and operate highly available infrastructure for an enterprise AI platform, spanning cloud systems, Kubernetes, automation, observability, and reliability engineering. Requires 5+ years of production infrastructure experience, strong Python or Go skills, and daily use of AI-assisted workflows.