Skip to content
PathAIPathAIBoston, MA

Senior/Staff Site Reliability Engineer - Data Center

This senior/staff SRE will design, operate, and secure data center infrastructure for machine learning workloads while integrating it with cloud systems. The role requires 8+ years of experience, production hardware and network expertise, automation and observability skills, and familiarity with hybrid infrastructure operations.

166k – 224k/yr
Remote8+ YOEDevOps / SRE

About the role

Responsibilities

  • Advance operations by implementing site reliability engineering best practices focused on users, monitoring, and automation.
  • Design, build, and operate data center infrastructure supporting a growing machine learning team.
  • Build highly secure on-premises environments that handle NIST and ISO standards.
  • Integrate on-premises data center environments with existing cloud infrastructure to create a seamless hybrid cloud environment.
  • Improve infrastructure reliability and resilience through root-cause analysis and reviews of design and implementation gaps.
  • Participate in platform on-call rotations and assist with urgent incident response.

Requirements

  • 8+ years of relevant experience.
  • Familiarity with modern data center network designs and comfort operating across network layers.
  • Experience administering physical hardware stacks in production, including iDRAC, IPMI, NVIDIA UFM, and Juniper systems.
  • Experience with virtualization, containerization, or container orchestration platforms such as EKS-Anywhere, ClusterAPI, and KVM.
  • Knowledge of storage solutions and optimization for high-performance workloads, including Quobyte, S3, FSx, and EFS.
  • Experience automating operational work through scripting and configuration management tools such as Ansible and Redfish.
  • Experience building monitoring infrastructure with observability tools such as Datadog, Grafana, and Prometheus.
  • Production operations experience, including critical infrastructure management, incident response, and scaling in rapidly growing environments.
  • Bachelor's degree in Computer Science or equivalent experience.
  • Intellectual curiosity and the ability to learn quickly in a complex environment.
  • Occasional travel to onsite data center locations.

Compensation

  • Annual pay range: $165,750–$224,450.
  • Not overtime eligible.

Skills

site reliability engineeringdata centershybrid cloudnetwork designidracipminvidia ufmjuniperkvmAnsibleredfishDatadogGrafanaPrometheusKubernetes

Similar roles

DevOps / SRE jobs
Crusoe

Staff Modular Data Center Engineer

CrusoeDenver, CO

Leads electrical design, optimization, and roadmap for prefabricated modular AI data centers (Crusoe Spark). Requires 5+ years in modular electrical systems, power distribution for AI compute, and cross-functional collaboration. In-office role in Denver with 10-20% travel.

168k – 192k/yrOn-site5+ YOEDevOps / SRE
Databricks

Staff Production Engineer- Public Sector

DatabricksVirginia

Owns secure cloud infrastructure, IAM, and automation across AWS, Azure, GCP for public sector environments. Requires 8+ years experience, deep cloud expertise, IaC tools like Terraform, and TS/SCI clearance eligibility.

162k – 223k/yrOn-site8+ YOEDevOps / SRE
Cerebras Systems

Member of Technical Staff (Software Engineer)

Cerebras SystemsSunnyvale, CA

Develops and optimizes Kubernetes-based infrastructure for high-performance AI inference services, including deployment, scaling, debugging, and integration with ML workflows. Requires Master's in CS and 1+ year experience with Docker, Kubernetes, Python, and related tools.

170k – 175k/yrRemoteDevOps / SRE
Shield AI

Senior Staff Network Architect (R4843)

Shield AIDallas, TX

Leads design, implementation, and optimization of complex network infrastructures with expertise in cybersecurity, hardware, data centers, and cloud/hybrid environments. Requires 10+ years experience, deep protocol knowledge, and hands-on enterprise networking skills.

170k – 250k/yrOn-site10+ YOEDevOps / SRE
Crusoe

Senior Staff Storage Systems Administrator

CrusoeSan Francisco, CA

Leads architecture, operation, and vendor strategy for petabyte-scale storage systems optimized for AI/HPC workloads in sustainable cloud infrastructure. Requires 10+ years experience with enterprise storage, scripting, and RFP/vendor management.

170k – 215k/yrOn-site10+ YOEDevOps / SRE