As a Principal Operations Engineer, Mechanical, you will be the senior technical authority for mechanical and cooling infrastructure across hyperscale AI data centers. You will lead site assessments, drive operational readiness, review designs, and ensure precision execution of critical systems.
150k – 250k/yrRemote10+ YOEDevOps / SRE
Principal Operations Engineer, Controls
FluidstackUnited States
As Principal Operations Engineer, Controls, you will be the senior technical authority for operational building automation and control systems across hyperscale AI data centers. You will lead site assessments, drive technical readiness, review designs, and ensure the continuous improvement of control systems.
150k – 250k/yrRemote10+ YOEDevOps / SRE
Principal Operations Engineer, Electrical
FluidstackUnited States
Fluidstack is seeking a Principal Operations Engineer, Electrical to be the senior technical authority for electrical infrastructure across their hyperscale AI data center portfolio. This role involves leading site assessments, driving technical readiness, reviewing designs, and feeding operational learnings back into the design and manufacturing organization.
150k – 250k/yrRemote10+ YOEDevOps / SRE
Software Engineer, GPU Infrastructure
FluidstackSan Francsisco, CA +3
Build and own automation, observability, and repair pipelines for one of the world's largest GPU compute fleets at hyperscale. Requires strong production engineering experience, hardware intuition at the firmware/silicon level, on-call ownership, and fluency with AI coding tools.
175k – 300k/yrOn-site5+ YOEDevOps / SRE
Software Engineer, Compute
FluidstackSan Francisco, CA +3
Build and own automation, observability, and repair pipelines for one of the world's largest GPU compute fleets. Requires hardware intuition at the firmware/silicon level, on-call ownership, and fluency with AI coding tools to eliminate toil at hyperscale.
208k – 269k/yrOn-site5+ YOEDevOps / SRE
Site Reliability Engineer, Compute
FluidstackSan Francisco, CA +3
Own end-to-end health, reliability, and automation of a massive GPU compute fleet for AI infrastructure. Build metrics, alerting, repair pipelines, GPU qualification platforms, and low-level BMC/Redfish tooling while driving incidents and using AI coding tools daily.
208k – 269k/yrOn-site5+ YOEDevOps / SRE
Network Deployment Tech Lead
FluidstackAustin, TX +3
Technical lead owning standards, tooling, automation, and escalation resolution for large-scale network deployments powering AI compute infrastructure. Requires senior experience leading clean network turn-ups, building ZTP/validation automation, and debugging from optics to routing.
186k – 220k/yrOn-site7+ YOEDevOps / SRE
Software Engineering, Commissioning Automation
FluidstackAustin, TX +3
Build and ship software that automates data center commissioning, including test orchestration, data capture, pass-fail analysis, and live reporting by integrating with BMS, EPMS, and test equipment. Requires experience building test automation or orchestration systems for hardware/infrastructure and replacing manual processes with trusted software.
269k – 317k/yrOn-site5+ YOEDevOps / SRE
Compute Deployment Engineer
FluidstackSan Francisco, CA +1
Own end-to-end compute deployment and rack qualification for large-scale GPU and accelerator fleets at Fluidstack, from facility handoff through burn-in, validation, and production readiness. Requires deep Linux/out-of-band management experience, hardware automation in Python/Go, data center operations, and methodical failure triage.
150k – 250k/yrHybrid5+ YOEDevOps / SRE
Principal Operations Engineer, Reliability
FluidstackAustin, TX
Own fleet reliability for nation-scale AI data centers: set availability targets, perform RCA on major incidents, build failure data pipelines, and define data-driven maintenance strategies. Requires hands-on experience improving critical infrastructure availability using Weibull, Pareto, and FMEA.
398k – 557k/yrOn-site7+ YOEDevOps / SRE
Facilities Production Technical Lead
FluidstackSan Francisco, CA
Lead technical direction for facilities production engineering, architecting telemetry pipelines from OT systems (BMS/EPMS/SCADA) into modern data stacks and setting controls integration standards across massive AI data center fleet.
188k – 266k/yrOn-site7+ YOEDevOps / SRE
Network Engineer, BMS/EPMS Networks
FluidstackNew York, NY
Own and design OT facility networks for BMS, EPMS, and controls traffic in AI data centers. Requires experience building industrial/OT networks, deep knowledge of controls protocols like BACnet/IP, and implementing segmentation that works with operations teams.
150k – 203k/yrOn-site5+ YOEDevOps / SRE
Distributed Systems Engineer
FluidstackSan Francisco, CA +3
Build and own the observability platform, production control plane, and fleet state as source of truth for a hyperscale GPU fleet powering AI infrastructure. Requires shipping scalable production services, on-call ownership, and comfort with AI coding tools; distributed systems and observability experience preferred.
208k – 269k/yrOn-site5+ YOEDevOps / SRE
Backend Engineer
FluidstackSan Francisco, CA +3
Production Engineer owning observability, control plane APIs, fleet state, and new hardware integration for hyperscale GPU infrastructure at an AI compute company. Requires strong automation mindset, on-call experience, and fluency with AI coding tools; distributed systems and observability stack experience preferred.
208k – 269k/yrOn-site5+ YOEDevOps / SRE
Software Engineer, Cloud Infrastructure
FluidstackSan Francsisco, CA +3
Build and own the observability platform, control plane APIs, and fleet state management for a hyperscale GPU infrastructure powering AI compute at 10-100s of GW scale. Requires production service ownership at scale, comfort with AI coding tools, and on-call incident response.
175k – 300k/yrOn-site5+ YOEDevOps / SRE
Platform Engineer
FluidstackSan Francisco, CA +3
Build foundational internal platforms (CMDB, DCIM, asset management, monitoring, automation) that power Fluidstack's global AI compute infrastructure and data center operations. Requires 3+ years building production systems with Python/Go, databases, IaC, and infrastructure/DevOps experience.
200k – 250k/yrOn-site3+ YOEDevOps / SRE
Production Engineer, Network
FluidstackAustin, TX
Own end-to-end network fleet health, monitoring, debugging tooling, and automated repair pipelines for massive AI datacenter infrastructure at Fluidstack. Requires systems thinking, automation-first mindset, on-call ownership, and daily use of AI coding tools like Claude/Cursor alongside Go/Python and network protocols.
175k – 300k/yrOn-site5+ YOEDevOps / SRE
Network Engineer, Design & Engineering
FluidstackNew York, NY +4
Design end-to-end datacenter network architectures for AI training and inference workloads. Own topology selection, fabric design, physical infrastructure integration, and produce deployable HLDs/LLDs across multiple GPU platforms and customer requirements.
180k – 300k/yrOn-site5+ YOEDevOps / SRE
Production Engineer, Compute
FluidstackSan Francisco, CA +3
Own end-to-end health, repair automation, and qualification of a hyperscale GPU/TPU compute fleet. Build metrics pipelines, firmware tooling, and self-healing repair workflows across Kubernetes and bare metal.
175k – 300k/yrHybrid5+ YOEDevOps / SRE
Infrastructure Deployment Engineer
FluidstackNew York, NY +3
Leads on-site deployment of data center physical infrastructure, managing contractors, performing QA/QC on fiber optics and cabling, and ensuring compliance with standards. Requires 5+ years experience, SME-level fiber optic expertise, bachelor's degree, and 40% travel readiness.
150k – 250k/yrOn-site5+ YOEDevOps / SRE
Network Engineer, Deployment & Integration
FluidstackNew York, NY +2
Hands-on network engineer deploying and validating large-scale AI datacenter fabrics, configuring switches, troubleshooting physical/optical layers, and coordinating cross-functional teams. Requires 3-7 years datacenter experience and 70-80% travel to onsite locations.
150k – 250k/yrOn-site3+ YOEDevOps / SRE
Search
Location
21 jobs
Job results
Principal Operations Engineer, Mechanical
FluidstackUnited States
As a Principal Operations Engineer, Mechanical, you will be the senior technical authority for mechanical and cooling infrastructure across hyperscale AI data centers. You will lead site assessments, drive operational readiness, review designs, and ensure precision execution of critical systems.
150k – 250k/yrRemote10+ YOEDevOps / SRE
Principal Operations Engineer, Controls
FluidstackUnited States
As Principal Operations Engineer, Controls, you will be the senior technical authority for operational building automation and control systems across hyperscale AI data centers. You will lead site assessments, drive technical readiness, review designs, and ensure the continuous improvement of control systems.
150k – 250k/yrRemote10+ YOEDevOps / SRE
Principal Operations Engineer, Electrical
FluidstackUnited States
Fluidstack is seeking a Principal Operations Engineer, Electrical to be the senior technical authority for electrical infrastructure across their hyperscale AI data center portfolio. This role involves leading site assessments, driving technical readiness, reviewing designs, and feeding operational learnings back into the design and manufacturing organization.
150k – 250k/yrRemote10+ YOEDevOps / SRE
Software Engineer, GPU Infrastructure
FluidstackSan Francsisco, CA +3
Build and own automation, observability, and repair pipelines for one of the world's largest GPU compute fleets at hyperscale. Requires strong production engineering experience, hardware intuition at the firmware/silicon level, on-call ownership, and fluency with AI coding tools.
175k – 300k/yrOn-site5+ YOEDevOps / SRE
Software Engineer, Compute
FluidstackSan Francisco, CA +3
Build and own automation, observability, and repair pipelines for one of the world's largest GPU compute fleets. Requires hardware intuition at the firmware/silicon level, on-call ownership, and fluency with AI coding tools to eliminate toil at hyperscale.
208k – 269k/yrOn-site5+ YOEDevOps / SRE
Site Reliability Engineer, Compute
FluidstackSan Francisco, CA +3
Own end-to-end health, reliability, and automation of a massive GPU compute fleet for AI infrastructure. Build metrics, alerting, repair pipelines, GPU qualification platforms, and low-level BMC/Redfish tooling while driving incidents and using AI coding tools daily.
208k – 269k/yrOn-site5+ YOEDevOps / SRE
Network Deployment Tech Lead
FluidstackAustin, TX +3
Technical lead owning standards, tooling, automation, and escalation resolution for large-scale network deployments powering AI compute infrastructure. Requires senior experience leading clean network turn-ups, building ZTP/validation automation, and debugging from optics to routing.
186k – 220k/yrOn-site7+ YOEDevOps / SRE
Software Engineering, Commissioning Automation
FluidstackAustin, TX +3
Build and ship software that automates data center commissioning, including test orchestration, data capture, pass-fail analysis, and live reporting by integrating with BMS, EPMS, and test equipment. Requires experience building test automation or orchestration systems for hardware/infrastructure and replacing manual processes with trusted software.
269k – 317k/yrOn-site5+ YOEDevOps / SRE
Get new-job notifications on iOS
Hotfix on iOS
Get a push summary when new jobs match your saved alerts.
Compute Deployment Engineer
FluidstackSan Francisco, CA +1
Own end-to-end compute deployment and rack qualification for large-scale GPU and accelerator fleets at Fluidstack, from facility handoff through burn-in, validation, and production readiness. Requires deep Linux/out-of-band management experience, hardware automation in Python/Go, data center operations, and methodical failure triage.
150k – 250k/yrHybrid5+ YOEDevOps / SRE
Principal Operations Engineer, Reliability
FluidstackAustin, TX
Own fleet reliability for nation-scale AI data centers: set availability targets, perform RCA on major incidents, build failure data pipelines, and define data-driven maintenance strategies. Requires hands-on experience improving critical infrastructure availability using Weibull, Pareto, and FMEA.
398k – 557k/yrOn-site7+ YOEDevOps / SRE
Facilities Production Technical Lead
FluidstackSan Francisco, CA
Lead technical direction for facilities production engineering, architecting telemetry pipelines from OT systems (BMS/EPMS/SCADA) into modern data stacks and setting controls integration standards across massive AI data center fleet.
188k – 266k/yrOn-site7+ YOEDevOps / SRE
Network Engineer, BMS/EPMS Networks
FluidstackNew York, NY
Own and design OT facility networks for BMS, EPMS, and controls traffic in AI data centers. Requires experience building industrial/OT networks, deep knowledge of controls protocols like BACnet/IP, and implementing segmentation that works with operations teams.
150k – 203k/yrOn-site5+ YOEDevOps / SRE
Distributed Systems Engineer
FluidstackSan Francisco, CA +3
Build and own the observability platform, production control plane, and fleet state as source of truth for a hyperscale GPU fleet powering AI infrastructure. Requires shipping scalable production services, on-call ownership, and comfort with AI coding tools; distributed systems and observability experience preferred.
208k – 269k/yrOn-site5+ YOEDevOps / SRE
Backend Engineer
FluidstackSan Francisco, CA +3
Production Engineer owning observability, control plane APIs, fleet state, and new hardware integration for hyperscale GPU infrastructure at an AI compute company. Requires strong automation mindset, on-call experience, and fluency with AI coding tools; distributed systems and observability stack experience preferred.
208k – 269k/yrOn-site5+ YOEDevOps / SRE
Software Engineer, Cloud Infrastructure
FluidstackSan Francsisco, CA +3
Build and own the observability platform, control plane APIs, and fleet state management for a hyperscale GPU infrastructure powering AI compute at 10-100s of GW scale. Requires production service ownership at scale, comfort with AI coding tools, and on-call incident response.
175k – 300k/yrOn-site5+ YOEDevOps / SRE
Platform Engineer
FluidstackSan Francisco, CA +3
Build foundational internal platforms (CMDB, DCIM, asset management, monitoring, automation) that power Fluidstack's global AI compute infrastructure and data center operations. Requires 3+ years building production systems with Python/Go, databases, IaC, and infrastructure/DevOps experience.
200k – 250k/yrOn-site3+ YOEDevOps / SRE
Production Engineer, Network
FluidstackAustin, TX
Own end-to-end network fleet health, monitoring, debugging tooling, and automated repair pipelines for massive AI datacenter infrastructure at Fluidstack. Requires systems thinking, automation-first mindset, on-call ownership, and daily use of AI coding tools like Claude/Cursor alongside Go/Python and network protocols.
175k – 300k/yrOn-site5+ YOEDevOps / SRE
Network Engineer, Design & Engineering
FluidstackNew York, NY +4
Design end-to-end datacenter network architectures for AI training and inference workloads. Own topology selection, fabric design, physical infrastructure integration, and produce deployable HLDs/LLDs across multiple GPU platforms and customer requirements.
180k – 300k/yrOn-site5+ YOEDevOps / SRE
Production Engineer, Compute
FluidstackSan Francisco, CA +3
Own end-to-end health, repair automation, and qualification of a hyperscale GPU/TPU compute fleet. Build metrics pipelines, firmware tooling, and self-healing repair workflows across Kubernetes and bare metal.
175k – 300k/yrHybrid5+ YOEDevOps / SRE
Infrastructure Deployment Engineer
FluidstackNew York, NY +3
Leads on-site deployment of data center physical infrastructure, managing contractors, performing QA/QC on fiber optics and cabling, and ensuring compliance with standards. Requires 5+ years experience, SME-level fiber optic expertise, bachelor's degree, and 40% travel readiness.
150k – 250k/yrOn-site5+ YOEDevOps / SRE
Network Engineer, Deployment & Integration
FluidstackNew York, NY +2
Hands-on network engineer deploying and validating large-scale AI datacenter fabrics, configuring switches, troubleshooting physical/optical layers, and coordinating cross-functional teams. Requires 3-7 years datacenter experience and 70-80% travel to onsite locations.