Senior/Staff Site Reliability Engineer - Data Center
This senior/staff SRE will design, operate, and secure data center infrastructure for machine learning workloads while integrating it with cloud systems. The role requires 8+ years of experience, production hardware and network expertise, automation and observability skills, and familiarity with hybrid infrastructure operations.
166k – 224k/yr
Remote8+ YOEDevOps / SRE
About the role
Responsibilities
Advance operations by implementing site reliability engineering best practices focused on users, monitoring, and automation.
Design, build, and operate data center infrastructure supporting a growing machine learning team.
Build highly secure on-premises environments that handle NIST and ISO standards.
Integrate on-premises data center environments with existing cloud infrastructure to create a seamless hybrid cloud environment.
Improve infrastructure reliability and resilience through root-cause analysis and reviews of design and implementation gaps.
Participate in platform on-call rotations and assist with urgent incident response.
Requirements
8+ years of relevant experience.
Familiarity with modern data center network designs and comfort operating across network layers.
Experience administering physical hardware stacks in production, including iDRAC, IPMI, NVIDIA UFM, and Juniper systems.
Experience with virtualization, containerization, or container orchestration platforms such as EKS-Anywhere, ClusterAPI, and KVM.
Knowledge of storage solutions and optimization for high-performance workloads, including Quobyte, S3, FSx, and EFS.
Experience automating operational work through scripting and configuration management tools such as Ansible and Redfish.
Experience building monitoring infrastructure with observability tools such as Datadog, Grafana, and Prometheus.
Production operations experience, including critical infrastructure management, incident response, and scaling in rapidly growing environments.
Bachelor's degree in Computer Science or equivalent experience.
Intellectual curiosity and the ability to learn quickly in a complex environment.
Occasional travel to onsite data center locations.
Compensation
Annual pay range: $165,750–$224,450.
Not overtime eligible.
Skills
site reliability engineeringdata centershybrid cloudnetwork designidracipminvidia ufmjuniperkvmAnsibleredfishDatadogGrafanaPrometheusKubernetes
Leads electrical design, optimization, and roadmap for prefabricated modular AI data centers (Crusoe Spark). Requires 5+ years in modular electrical systems, power distribution for AI compute, and cross-functional collaboration. In-office role in Denver with 10-20% travel.
168k – 192k/yrOn-site5+ YOEDevOps / SRE
Staff Production Engineer- Public Sector
DatabricksVirginia
Owns secure cloud infrastructure, IAM, and automation across AWS, Azure, GCP for public sector environments. Requires 8+ years experience, deep cloud expertise, IaC tools like Terraform, and TS/SCI clearance eligibility.
162k – 223k/yrOn-site8+ YOEDevOps / SRE
Member of Technical Staff (Software Engineer)
Cerebras SystemsSunnyvale, CA
Develops and optimizes Kubernetes-based infrastructure for high-performance AI inference services, including deployment, scaling, debugging, and integration with ML workflows. Requires Master's in CS and 1+ year experience with Docker, Kubernetes, Python, and related tools.
170k – 175k/yrRemoteDevOps / SRE
Senior Staff Network Architect (R4843)
Shield AIDallas, TX
Leads design, implementation, and optimization of complex network infrastructures with expertise in cybersecurity, hardware, data centers, and cloud/hybrid environments. Requires 10+ years experience, deep protocol knowledge, and hands-on enterprise networking skills.
170k – 250k/yrOn-site10+ YOEDevOps / SRE
Senior Staff Storage Systems Administrator
CrusoeSan Francisco, CA
Leads architecture, operation, and vendor strategy for petabyte-scale storage systems optimized for AI/HPC workloads in sustainable cloud infrastructure. Requires 10+ years experience with enterprise storage, scripting, and RFP/vendor management.