Latest Cloud Infrastructure jobs
Job results
Own end-to-end supplier integration, performance management, and system connectivity to synchronize supplier capacity, quality, and delivery with manufacturing build plans. Requires prior experience managing supplier launches, recovering failing suppliers, and connecting processes to ERP/MRP systems.
Lead the TPM function for R&D at Fluidstack, running portfolios of hardware and infrastructure programs across mechanical, electrical, controls and modelling workstreams. Own executive reporting, build the TPM team, and drive programs from modular designs to scheduled build reality. Requires prior leadership of hardware/infrastructure TPM teams and portfolio management experience.
Lead the architecture and implementation of facilities software automation for physical plant control, from telemetry to actuation with built-in safety. Own closed-loop control, simulation, rollout across sites, and set engineering standards while continuing to write critical code. Requires prior leadership of physical systems automation teams.
Coordinate tenant improvement construction, track against production ramp, manage punch/turnover and construction paperwork (RFIs, submittals, change orders) for AI data center factories. Requires commercial/industrial construction coordination experience from owner/GC side and managing TI in operating facilities.
Design and defend network security architecture for AI compute infrastructure at gigawatt scale across data centers, OT networks, and cloud. Build detection, harden OT/BMS systems, and automate controls using IaC. Requires scaling network security experience, packet analysis, IT/OT segmentation, and real intrusion detection.
Own compute and storage sourcing (GPU servers, storage systems, node components) from supplier strategy through delivery for massive AI data center buildout. Negotiate pricing/lead times/capacity with OEMs/ODMs, dual-source critical configs, and manage supplier performance tied to forward pipeline.
Own the structural steel category for Fluidstack's massive AI compute buildout. Source and qualify fabricators at industrial scale, negotiate commodity-resistant pricing, balance capacity across shops, and manage hedging to keep modular data center construction off the critical path.
Build and own ML/LLM systems for internal operations including forecasting, risk flagging, and document extraction. Ship production agentic systems end-to-end with guardrails and partner with data engineering to integrate predictions into tools.
Own end-to-end sourcing and supplier qualification for mechanical categories (steel, fabrications, cooling components) to support massive AI data center manufacturing scale-up. Requires deep manufacturer-side sourcing experience, ability to read drawings/specs, secure capacity, and negotiate long-term agreements.
Design and engineer electrical distribution systems for AI data centers from medium voltage to the rack. Run power studies in ETAP/SKM, develop equipment specifications, and support field builds and energization.
Build and integrate software for the SCADA stack at data centers powering frontier AI compute, including services, EPMS integrations, historians, alarms, and repeatable deployment tooling. Requires experience with industrial control systems, OPC UA, and production-grade software for mission-critical environments.
Lead robotics deployment projects into live data center and factory sites, coordinating vendors, integrators, and teams from pilot through scale-up while owning the deployment playbook, safety, commissioning, and post-go-live performance tracking.
Engineer automated commissioning for data center control systems by embedding test sequences into BMS/PLC logic, building test libraries, executing on live sites, analyzing failures, and driving upstream design fixes. Requires hands-on commissioning experience writing test procedures for building/data center controls.
Develop and qualify suppliers for AI data center manufacturing through hands-on process audits, PPAP approvals, root cause quality fixes, and capacity building at supplier sites. Requires proven manufacturing supplier development experience with rigorous on-site problem solving.
Build software services, APIs, and generative tooling for programmable infrastructure modeling in CAD/BIM platforms. Requires experience with engineering design data, automation replacing manual work, geometric/graph data, and production code integrating vendor APIs.
Own reliability engineering for AI data center reference designs. Build RAM models, run cross-discipline FMEAs, analyze field failures, and drive design changes for high-availability infrastructure.
Model and reconcile IT/hardware capacity needs across data center sites, building automated views for planning and procurement while chasing discrepancies in orders, inventory, and lead times.
Design and engineer mechanical utilities including fuel gas, water, compressed air, and balance-of-plant systems for behind-the-meter power generation plants supporting AI compute infrastructure. Requires experience with utility systems for power generation or heavy industry, fuel gas design, and EPC package review.
Build and deploy demand management control strategies that flex gigawatts of data center load in response to grid signals and prices while coordinating IT, cooling, and generation assets without violating SLAs. Requires prior experience engineering controls for energy systems such as demand response or microgrids, power systems knowledge, utility/market integration, and simulation-first validation.
Senior Manager responsible for owning end-to-end finance systems integrations and implementations across NetSuite, Pigment, banks and operational systems at a high-growth AI compute infrastructure company. Requires proven experience delivering automated financial data flows and managing vendors for accounting-close reliability.
Project manager responsible for deploying controls and sensor systems (BMS/EPMS) across data center sites, coordinating installation, vendors, construction, and commissioning to close projects on time with zero punch list items. Requires prior experience managing controls/instrumentation installations on industrial or data center projects.
Own and engineer rotating equipment (turbines, engines, generators) for behind-the-meter power generation at AI data centers. Specify, integrate, test, and support reliability of mechanical packages including vibration and alignment.
Own the electrical equipment category (switchgear, transformers, UPS, generators, busway) for massive AI compute buildout. Negotiate multi-year capacity reservations with OEMs, dual-source non-interchangeable equipment, track factory production, and expedite as needed for data center construction.
Senior Named Account Executive responsible for driving multi-million dollar enterprise sales of Cloudflare's networking, security, and edge computing platform. Requires 8+ years of B2B complex tech sales experience, deep technical understanding of customer IT architectures, virtual team leadership, and consistent quota overachievement.
Escalation Engineer (Tier 3) who owns the hardest customer support issues for Nerdio's Microsoft cloud products (AVD, Intune, Azure). Performs deep root-cause analysis, collaborates with Product/Engineering, mentors T1/T2 engineers, creates RCAs and KB articles, and drives systemic improvements. Requires 4+ years technical support experience and strong Microsoft cloud troubleshooting skills.
Senior GTM Specialist driving Cloudflare One (SASE platform) revenue growth and adoption. Advise sales leaders, enable AEs on positioning, support strategic deals, build partner relationships, and provide feedback to product teams. Requires 7+ years in cybersecurity/SaaS sales with deep technical expertise in SASE/SSE solutions.
Build and operate highly available, low-latency microservices and APIs for Kong Konnect’s cloud-hosted control plane. The role requires 3+ years of distributed software experience, strong Go and database skills, and experience with Kubernetes, observability, and production SaaS reliability.
Build and operate distributed systems infrastructure for data localization and geographic compliance on Cloudflare's global edge network. Requires 3+ years building production distributed systems with deep knowledge of consistency models, failure modes, and languages like Go or Rust.
Provides technical support for Crusoe Cloud’s GPU and HPC infrastructure, handling customer issues, incident triage, troubleshooting, and on-call response. Requires strong Linux and cloud-platform skills, Kubernetes or workload-management experience, and at least five years in customer support.
Build and own end-to-end security for Fluidstack's bare metal AI compute fleet, from supply chain to decommissioning. Harden Linux, enforce BMC security, implement zero-trust networking and encryption at gigawatt scale.
Own end-to-end compute deployment and rack qualification for large-scale GPU and accelerator fleets at Fluidstack, from facility handoff through burn-in, validation, and production readiness. Requires deep Linux/out-of-band management experience, hardware automation in Python/Go, data center operations, and methodical failure triage.
Own global cash and liquidity operations for an AI infrastructure company. Build the treasury function from the ground up, manage banking relationships, cash flow forecasting, FX risk, and multi-currency operations in a high-growth environment.
Serve as People Partner for client organizations at a fast-scaling AI compute infrastructure company. Handle employee relations, coach first-time managers, run people processes at scale, and leverage AI to automate HR operations.
Serve as on-site People Partner for Fluidstack's rapidly scaling Manufacturing organization. Coach first-time managers, own employee relations and investigations, and build HR rhythms for hourly and salaried factory workforce.
Own and optimize talent systems including ATS configuration, automation for scheduling/scoring, hiring dashboards, and end-to-end process changes to enable high-velocity recruiting at an AI infrastructure company.
Build and maintain the software layer for data center facility controls including BMS/EPMS integrations, alarm pipelines, rationalization, and automated responses. Requires experience writing software against industrial controls protocols (BACnet, Modbus, OPC-UA) and fluency in Python or Go.
Lead the facilities production engineering team at Fluidstack, owning hiring, roadmap, delivery, and on-call health while bridging OT/IT skills for large-scale AI data center operations. Requires experience managing engineering teams in operational environments and balancing roadmap vs. interrupt-driven support.
Lead technical direction for facilities production engineering, architecting telemetry pipelines from OT systems (BMS/EPMS/SCADA) into modern data stacks and setting controls integration standards across massive AI data center fleet.
Own reliability, SLAs, and escalations for customer AI/HPC workloads at massive scale. Debug full-stack issues (hardware to scheduler), deliver technical customer incident communications, and drive root-cause fixes with internal engineering teams.
Own end-to-end employee relations for a fast-growing AI infrastructure company, including investigations, conflict resolution, manager coaching, and identifying systemic issues. Requires prior ER ownership at a scaling organization with legally sound documentation practices.
Lead sourcing and procurement for site readiness services including power, water, security, waste, and facilities to enable rapid data center deployment. Build playbooks, negotiate multi-site agreements, and coordinate with construction/operations teams for greenfield industrial sites.
Serve as HR business partner to data center operations leadership, handling org design, performance management, shift workforce scheduling, retention, and end-to-end employee relations for hourly and salaried staff in 24/7 onsite environments.
Lead the build-out of the People function at a hyper-growth AI infrastructure company, owning end-to-end employee lifecycle, AI-powered operations, culture, and multi-jurisdictional compliance to support 10x headcount scaling.
Own and maintain onsite construction schedules in Primavera P6 for gigawatt-scale data centers. Build crew-logic schedules, run weekly lookaheads, analyze risks, and align with commissioning to drive priorities on multiple simultaneous builds.
Lead construction operations for multiple concurrent gigawatt-scale data center sites. Own schedule, cost, quality, and safety; standardize playbooks; drive constructability feedback into design. Requires experience running large industrial/mission-critical projects and holding GCs to aggressive schedules.
Mechanical Engineer owning cooling and piping scope for rapid gigawatt-scale data center builds. Review designs against reference, answer field RFIs quickly, and support construction/commissioning. Requires data center or industrial plant mechanical engineering experience.
Lead electrical engineering for high-speed data center builds delivering gigawatts of AI compute capacity. Own reference designs, field decisions, energization, and critical review of studies for utility-to-rack power distribution.
Own controls engineering (BMS/EPMS) for gigawatt-scale AI data centers: design reviews, integrator oversight, and rapid commissioning support on mission-critical projects built in months. Requires experience engineering building/industrial controls, reading/writing sequences of operations, reviewing integrator work, and debugging from points lists.
Own supplier quality for critical electrical equipment (switchgear, transformers, UPS, busway) at Fluidstack. Conduct factory audits, witness testing, set contractual quality requirements, and drive corrective actions to support rapid 10-100GW AI compute buildout.
Own supplier quality for mechanical equipment (CDUs, chillers, pumps, piping) from qualification through delivery at a company building massive AI compute infrastructure. Conduct audits, set inspection/test plans, witness FATs, and drive on-floor root cause/corrective actions.