Latest Cloud Infrastructure jobs
Job results
Customer Engineer embedded in GTM for Cloudflare's Digital Native Enterprise accounts. Own technical relationships across the full customer lifecycle: pre-sales discovery/PoCs, onboarding, adoption, expansion with quota, and C-level advisory on security, networking, and developer platforms.
Senior Customer Engineer serving as a quota-carrying technical advisor at Cloudflare. Owns full customer lifecycle from technical discovery, PoCs, and architecture design through onboarding, adoption, QBRs, and expansion for Security, Networking, and Developer Platforms.
Staff Product Manager owning vision and roadmap for Render's observability, spanning systems telemetry, OpenTelemetry, agent/LLM traces, and AI cost tracking. Requires 7+ years PM experience in observability/infra/devtools, deep OTel fluency, and AI workload economics knowledge.
Staff Product Manager owning vision and roadmap for Render's managed datastores with Postgres at the center. Define modern database experiences for AI-native and agentic workloads, drive reliability, scaling, and developer experience in close partnership with engineering.
Build and optimize the end-to-end LLM inference stack for production deployments. Profile and tune serving frameworks (vLLM/SGLang) and CUDA kernels for latency, throughput and cost; partner directly with customer teams to take workloads from POC to monitored production.
Lead executive communications for Cloudflare's co-founder and President, including media, speaking, and branding. Build and scale company-wide executive comms programs, translate deep-tech topics for diverse audiences, and integrate AI tools to optimize workflows. Requires 10+ years in executive comms for C-level leaders.
Sales Development Representative responsible for researching enterprise accounts, multi-channel outbound prospecting, and scheduling meetings with the Field Sales team to drive interest in Nasuni's hybrid cloud file data platform. Requires passion for B2B sales, strong communication, and tech enthusiasm.
Build and operate distributed backend services enforcing geographic data residency boundaries on Cloudflare's global edge network. Requires 5+ years building production distributed systems with deep knowledge of consistency, failure modes, and compliance constraints; full-stack ownership in Go/Rust.
Build and deploy AI-powered internal tools and shared infrastructure to automate workflows and boost productivity across engineering, product, operations, and business teams at a frontier AI compute infrastructure company. Requires 3+ years production software engineering, full-stack skills, and deep hands-on experience with LLMs, agents, and AI coding tools.
Own mechanical operations and escalation for chiller plants, CRAC/CRAH units, and cooling systems in live gigawatt-scale AI data centers. Lead preventive maintenance, root-cause analysis, procedure development, and commissioning handovers while working rotating shifts in a fast-scaling environment.
Lead site procurement strategy for large-scale US data center development. Source, evaluate, and secure powered land (200MW+), conduct due diligence on power/buildability/permitting, negotiate leases, and manage stakeholder relationships for hyperscale AI infrastructure.
Lead the compute production engineering team at Fluidstack, owning availability SLOs, automation for node lifecycle (provisioning to remediation), and on-call models for tens of thousands of GPUs at nation-scale. Requires prior leadership of large-fleet SRE/production teams, proven availability improvements, and automated remediation experience.
Lead the IaaS team at Fluidstack to deliver high-scale compute infrastructure for AI workloads. Own the platform layer for provisioning, isolation, networking, storage, and self-service capabilities while setting SLOs and balancing roadmap with operations.
Own environmental and civil engineering for multi-GW data center sites, leading stormwater, grading, drainage, and permitting (SWPPP, wetlands) from site screening through approvals while coordinating with design teams and managing consultants to maintain aggressive construction schedules.
Produce and maintain civil drawing sets for large-scale data center site development using Civil 3D, from redlines to issued-for-construction. Coordinate with other disciplines, maintain CAD standards, and deliver rapid revisions to meet tight permitting and construction deadlines.
Manage end-to-end workforce housing program for Crusoe's construction sites, including capacity planning tied to labor schedules, site selection, vendor management, operations oversight with hospitality standards, and financial modeling. Requires experience standing up complex multi-stakeholder programs and leading field teams.
Senior Project Manager leading planning and execution of Engineered-to-Order modular data center building projects from concept through delivery. Requires 7+ years PM experience in construction/engineering/manufacturing, strong change/risk/budget management skills, and BS degree.
Lead the full lifecycle development of phased array antenna systems for satellite ground stations, from architecture and requirements through design, integration, testing, production, and deployment. Requires 7+ years experience owning complex multidisciplinary hardware products with deep expertise in at least one of mechanical, electrical, RF, or embedded engineering.
Senior Performance Engineer responsible for Linux kernel optimization, system benchmarking, and low-level performance tuning to enhance Crusoe's AI cloud infrastructure. Requires deep Linux kernel expertise, proficiency in Go/C/C++, and hands-on experience with performance optimization in complex environments.
Own end-to-end NPI and R&D programs for next-gen data center hardware, coordinating mechanical/electrical/controls engineering, manufacturing, supply chain, and deployment teams from concept to production at multi-GW scale. Requires 5+ years TPM experience in build-heavy hardware or infrastructure domains with strong bias to action.
Manager, Revenue Accounting responsible for end-to-end ownership of complex revenue streams, technical accounting under ASC 606/842, contract analysis, operationalizing processes, month-end close, and audit support in a fast-growing AI infrastructure company. Requires CPA, 5+ years experience, deep GAAP expertise, and self-starter mindset.
Lead a team of 4-6 engineers building and operating Crusoe's telemetry agent for metrics/logs from hosts and GPUs. Own delivery of next-gen agent releases against fixed deadlines while hiring, coaching staff-level engineers, and maintaining high operational standards for fleet-wide software.
Legal Counsel - Corporate will lead high-velocity debt financings, capital markets, and strategic transactions for Crusoe's AI infrastructure business. Requires 7-10 years of complex transactional experience, JD, and strong business judgment to balance legal risks with commercial objectives.
Build and deploy internal AI-powered tools and shared infrastructure (agents, skills, LLM integrations) to automate workflows and boost productivity across engineering, operations, and business teams at an AI compute infrastructure company. Requires 3+ years full-stack experience, strong LLM/agent expertise, and autonomous shipping with great product taste.
Build and own automation, observability, and repair pipelines for one of the world's largest GPU compute fleets at hyperscale. Requires strong production engineering experience, hardware intuition at the firmware/silicon level, on-call ownership, and fluency with AI coding tools.
Build and own automation, observability, and repair pipelines for one of the world's largest GPU compute fleets. Requires hardware intuition at the firmware/silicon level, on-call ownership, and fluency with AI coding tools to eliminate toil at hyperscale.
Own end-to-end health, reliability, and automation of a massive GPU compute fleet for AI infrastructure. Build metrics, alerting, repair pipelines, GPU qualification platforms, and low-level BMC/Redfish tooling while driving incidents and using AI coding tools daily.
Model and analyze capacity (power, space, cooling, compute) across AI data center fleet. Track consumption, build executive reporting, and run scenario analysis to inform allocation and deal decisions. Requires infrastructure capacity analysis experience and SQL/Python modeling skills.
Lead physical security operations and teams across multiple data center sites in a region. Own end-to-end posture, standardize procedures, support customer audits, and integrate security into facility growth for frontier AI compute infrastructure.
Director of Accounting to own and scale Fluidstack's global payroll operations end-to-end. Build processes and systems for a rapidly growing, multi-jurisdictional workforce while ensuring compliance, accuracy, and data integrity in partnership with Finance, Legal, and HR.
Own the data center lab floor at Fluidstack Labs by performing rack-and-stack, cabling, hardware troubleshooting, inventory, RMAs, and smart-hands support for first-sample AI compute, networking, and liquid-cooled systems. Requires deep server/network hardware experience, Linux console skills, and meticulous asset documentation.
Leads the founding Tel Aviv Production Engineering team, combining people management with hands-on reliability engineering, incident response, automation, and firmware optimization. Requires 8+ years in infrastructure, SRE, or production engineering and 2+ years of direct engineering leadership.
Own end-to-end logistics and asset lifecycle management for hyperscale data centers, including receiving, shipping, inventory accuracy at 99%+, parts fulfillment, RMAs, and vendor coordination in a fast-paced construction environment.
Own the cost model and produce estimates for Fluidstack's modular data center product across structural, mechanical, electrical, and controls scopes. Provide rapid pricing for design trade-offs, interrogate vendor quotes, and reconcile estimates against actuals to improve accuracy.
Technical Program Manager responsible for executing Fluidstack's security program across physical, logical, and datacenter domains. Drive cross-functional deliverables, compliance initiatives, risk tracking, and operating rhythms to secure frontier AI compute infrastructure.
Welding Engineer responsible for qualifying WPS/PQR to AWS D1.1, inspecting/certifying welds, building weld quality programs, qualifying welders, and reducing defects for high-rate modular data center manufacturing. Requires current CWI credential and hands-on experience writing procedures and running quality programs in production environments.
Own master schedule, cost controls, and risk register for a portfolio of gigawatt-scale data center construction programs. Build executive reporting, program controls standards, and coordinate across construction, procurement, and design teams.
Own monthly/quarterly financial close, build technical accounting policies for massive data center capex, stand up reporting systems from the ground up, run clean external audits, and compress close cycles at a fast-scaling AI compute infrastructure company.
Build and maintain demand, lead-time, and capacity forecasts for critical data center equipment to keep multi-GW AI compute build pipelines on schedule. Partner with procurement using data-driven models and dashboards to mitigate risks and optimize multi-hundred-million-dollar commitments.
Own end-to-end ASC 740 income tax provisions for a rapidly scaling multi-jurisdictional data center company. Build processes/controls from scratch, manage external advisors/auditors, model tax impacts of expansion, and partner with accounting/FP&A on closes and disclosures.
Source, diligence, negotiate, and secure powered land and brownfield sites delivering hundreds of MW to GWs of capacity for AI data centers. Requires hands-on experience closing power-constrained land deals, reading interconnection queues, running full diligence, and managing pipelines while moving faster than counterparties.
Senior Manager owning technical accounting memos for complex data center transactions (revenue, leases, equity), running financial close, building policies/controls from scratch, and managing external audits as the company scales toward public readiness. Requires experience with novel transactions, GAAP reporting, and audit scrutiny.
Lead end-to-end incident response for frontier AI infrastructure, building detection logic, playbooks, and forensic tooling from the ground up while investigating threats across cloud, endpoint, network, and physical systems at gigawatt scale. Requires hands-on experience leading major incidents, proactive threat hunting, and standing up IR programs.
Build and automate on-site IT infrastructure (networks, identity, endpoints) for gigawatt-scale AI data center campuses. Own multi-tenant security, repeatable deployments, and rapid incident response for fast-growing field/construction teams.
Own manufacturing engineering for heavy chassis fabrication and weldment production at Fluidstack, developing process plans, launching lines, running PFMEAs, and driving yield/cycle time improvements for AI data center infrastructure.
Lead the hardware qualification lab for next-gen compute, network, and rack systems ahead of large-scale AI infrastructure deployment. Own qualification roadmap, lab operations, vendor relationships, and final standardization decisions backed by test data.
Lead mechanical engineering R&D for cooling, piping, and mechanical systems in modular data centers and power infrastructure. Own reference designs, enforce technical standards, review vendor packages, and grow the team while solving the hardest problems.
Mechanical Engineer responsible for designing and modeling demand-flexible cooling systems for AI data centers, including thermal inertia modeling, field validation, and turning cooling infrastructure into a grid asset. Requires experience with dynamic thermal modeling and HVAC systems for data centers or industrial plants.
Senior Manager owning global indirect tax (VAT/GST/sales & use) for a fast-growing AI data center company. Build compliance, registrations, filings, reclaims, and advise on structuring from first principles across multiple countries.
Fire Protection Engineer responsible for designing detection, suppression, and alarm systems for AI data center modular units and plants. Owns hydraulic calculations, layouts, standardized fabrication designs, permitting, and commissioning for mission-critical facilities.