Latest Cloud Infrastructure jobs
Job results
BIM Engineer responsible for model management, clash detection, and producing field deliverables for high-speed data center builds. Requires fast production skills in Revit and Navisworks plus experience on large MEP-heavy projects.
Own fire protection engineering for data center sites, including design reviews of sprinklers and detection systems, field support, AHJ coordination, and fast RFI/code question resolution. Requires commercial/industrial FPE experience, hydraulic calc review, and AHJ testing background.
Lead the end-to-end compute deployment process for large-scale GPU clusters at data centers, owning schedules, building scalable automation and workflows, managing mixed engineering/technician teams, and enforcing strict validation standards to rapidly turn over halls into production compute.
Own the full hardware fleet lifecycle for massive AI compute infrastructure, including receipt, deployment, spares strategy, RMA, tracking, and decommissioning across data center sites. Requires experience managing tens of thousands of hardware assets with strong focus on spares optimization, audit-ready records, and warranty recovery.
Lead TPM responsible for running multi-site data center deployment programs at massive scale, owning schedules, cross-site sequencing decisions, accurate leadership reporting, and building the TPM team to deliver GWs of AI compute infrastructure rapidly.
Lead mechanical engineering for Fluidstack's gigawatt-scale AI data centers, owning reference design, field decisions, and execution for direct-to-chip liquid cooling systems from concept to energized. Requires prior leadership of data center or heavy industrial mechanical programs with deep liquid cooling expertise.
Lead design management for Fluidstack's multi-gigawatt AI data center builds. Set and enforce aggressive design schedules across multiple firms, run high-velocity review processes for constructability/cost/reference conformance, and close the design-field loop on RFIs to enable months-not-years delivery at 50GW+ scale.
Own electrical engineering for data center sites including design reviews, field RFI support, and energization readiness for gigawatt-scale AI compute infrastructure. Requires data center or industrial electrical experience, fluency with one-lines/studies, and rigorous pre-energization processes.
Operate and evolve Render's scalable Key/Value (KV) datastore product at fleet scale. Take end-to-end ownership of full-stack projects, collaborate with product/design, participate in on-call, and drive new capabilities for millions of users.
Product Engineer owning end-to-end development and operations for Render's scalable Key/Value (Valkey) datastore platform. Requires 6+ years shipping software, experience operating KV stores at scale, and comfort with distributed systems.
Operate and evolve Render's managed Postgres fleet at scale. Take end-to-end ownership of full-stack projects to improve the Postgres product for millions of developers, with on-call duties.
Own end-to-end accounts payable processing including invoice matching, vendor management, and weekly payments while also supporting AR invoicing and collections in a fast-paced space infrastructure startup. Requires 1-4 years accounting experience, GAAP knowledge, ERP proficiency, and ability to obtain Top Secret clearance.
Own the full sales cycle for commercial accounts, from outbound prospecting and discovery through close, while helping shape an early-stage sales motion. The role requires 3–6 years of B2B SaaS closing experience, quota attainment, and comfort selling to both technical and finance or operations leaders.
Lead the end-to-end threat intelligence production pipeline at Cloudflare, managing schedules, quality gates, editing, and publication of reports and blogs. Requires 7+ years combined experience in technical writing, intelligence analysis, and cybersecurity, plus leadership in editorial/production functions.
Develops high-performance Go backend systems, SDKs, CLI tools, declarative state managers, and Terraform integrations for Kong’s API management platform. Requires 5+ years of backend engineering experience, strong API design skills, and production experience with infrastructure-as-code or platform tooling.
REACT Consultant performing incident response, active edge mitigation, forensics, and threat containment for Cloudflare customers across on-prem, cloud, and hybrid environments. Requires 3+ years cybersecurity experience including 2+ years in IR/forensics, strong network and OS knowledge, and customer-facing skills.
Customer Experience Manager responsible for post-sale customer lifecycle, retention, renewals, upsells, and building strategic relationships as a trusted advisor. Requires 3+ years in Customer Success or sales, fluency in Russian, and AI competencies for optimization and automation.
Design and manage sales compensation plans for all customer-facing teams at Cloudflare, architecting AI-native frameworks and new metrics for developer/AI monetization. Lead C-level discussions with CRO/CFO/CPO, own executive steering committee, and drive 3-year roadmap for compensation strategy.
Senior Mechanical Design Engineer responsible for designing structures, mechanisms, tooling, and RF waveguide components for global satellite ground station antennas. Requires 5+ years hardware design experience from concept to production, strong mechanical/thermal principles knowledge, and on-site work in Los Angeles.
Seasoned Product Manager to own Quote to Cash products and processes at Cloudflare, driving billing, invoicing, pricing, ERP, and AI monetization initiatives. Requires proven billing/FinTech experience at SaaS companies, strong cross-functional leadership, and ability to balance multiple timescales and domains.
Software Engineer building full-stack features for a gaming payments and e-commerce platform. Requires 1+ years experience with React and TypeScript; must work onsite in NYC.
Senior full-stack engineer building and scaling a payments and e-commerce platform for game publishers. Requires 5+ years experience with strong React and TypeScript skills, plus experience leading projects.
Lead on-site security posture for high-value AI compute infrastructure at data centers. Own all security systems, policies, vendor accountability, incident response, and represent security in site planning and operations.
Build and ship production AI agents on Cloudflare's edge platform using Workers, Durable Objects, and AI tools. Requires strong TypeScript/Rust experience, observability expertise, and hands-on LLM tooling for evals, safety, and multi-agent systems.
Lead technical investigations, threat hunting, detection development, and response for insider threats. Partner closely with Legal, HR, and Privacy teams while ensuring compliance with regulatory, legal, and ethical standards. Requires 5+ years in security with forensics/investigations focus.
12-14 week full-time Software Engineer internship at Cloudflare. Work on projects impacting millions of internet users, ship code to production, collaborate with mentors and cross-functional teams, and present work company-wide. Requires pursuit of CS/Engineering degree, curiosity, and ability to work onsite in Austin 3-5 days/week.
Serve as the senior technical operations expert for Crusoe's SiteOps organization, owning Tier 2/3 GPU hardware escalations, OEM/ODM partnerships, platform standards, SOPs, training, and cross-functional collaboration at HQ while traveling to sites as needed. Requires 7+ years of hands-on GPU data center operations experience.
Senior Digital Marketing Manager owning strategy and execution of paid media, email, webinars, PLG, and ABM campaigns to drive lead generation, pipeline, and brand awareness at an AI infrastructure company. Requires 7+ years B2B digital marketing experience with deep expertise in Google/LinkedIn/Meta ads, Marketo, Salesforce, and analytics.
Tax Analyst responsible for preparing income tax provisions, managing US tax filings, assisting with M&A due diligence, stock compensation analysis, R&D credits, and driving AI/automation-based digital transformation of the tax function. Requires 2+ years tax experience, bachelor's in accounting/finance, Excel and AI proficiency.
Lead and grow teams of product engineers at Render, partnering with PMs and designers to translate user needs into scalable developer platform features. Requires 8+ years engineering experience and 4+ years managing teams with strong technical and product mindset.
Lead Cloudflare's digital communications strategy, overseeing social media, executive positioning, and multi-platform storytelling to build brand influence with developers, enterprises, and tech leaders. Requires 10+ years in communications with deep tech understanding and leadership of digital teams.
Senior Data Engineer building scalable data pipelines, services, and AI-ready data layers in Go/Scala/ClickHouse to power internal products, analytics, and agentic AI for go-to-market, engineering, and product teams. Requires 5+ years experience in production data systems, strong programming, SQL, and databases.
Senior Product Designer crafting complex web UIs for a developer cloud platform. Own observability, multi-service dashboards, deployments, error states, and first-run experiences with obsessive craft and information architecture.
Staff Product Designer owning complex multi-service observability, deployments, and failure states for Render's developer cloud platform. Requires 7+ years experience, exceptional craft, information architecture skills, and fluency in developer workflows to create cohesive, high-stakes interfaces.
Senior Software Engineer owning Crusoe Cloud's managed container registry. Build, scale, and optimize core distributed services for AI training/inference workloads; drive performance, reliability, and architecture decisions in production.
Staff Production Engineer responsible for developing automation/observability, scaling virtualization (KVM/QEMU), optimizing Linux kernel performance, and supporting AI/HPC workloads on CPU/GPU/DPU hardware. Requires 8+ years in Linux systems engineering, kernel internals, and virtualization.
The Enterprise Sales Development Representative will generate and qualify inbound and outbound opportunities for Azure Virtual Desktop solutions, partnering with Microsoft sales teams, channel partners, and direct customers. The role requires 1–3 years of sales or business development experience, cloud technology familiarity, and English plus German, French, or Dutch fluency.
Build and scale the control plane for Cloudflare Workers, the serverless edge platform. Design distributed systems, high-traffic APIs, storage schemas, and own production reliability for services powering Pages and R2. Requires strong Go experience plus knowledge of observability, databases, and Kubernetes.
Lead architecture and strategy for Cloudflare's Workers control plane as the most senior engineer on the Deploy & Config team. Own distributed systems, high-scale APIs, reliability, and developer experience powering the serverless edge platform.
Lead day-to-day operations of Crusoe's 24/7 Global Security Operations Center (GSOC), overseeing monitoring, incident response, intelligence, and crisis management to protect facilities and infrastructure. Requires experience building high-performing teams, driving operational maturity, and ensuring compliance in a fast-growing environment.
Build and optimize Salesforce and the modern GTM technology ecosystem (HubSpot, Workato, Snowflake, etc.) as the core of a composable, event-driven, agentic revenue platform. Requires 7+ years enterprise engineering experience including 5+ years hands-on Salesforce development.
Chief of Staff to the CTO at Northwood Space, acting as a strategic partner and force multiplier to drive execution of complex engineering and product initiatives. Requires 5+ years in technical program management or similar, strong organizational skills, and a bachelor's in engineering.
Legal Counsel supporting Crusoe's corporate finance and strategic transactions, including debt financings, capital markets deals, drafting/negotiation of credit agreements, due diligence, and covenant compliance. Requires JD, 7-10 years of complex transactional experience in finance/corporate matters, strong business judgment, and ability to manage high-stakes deals in a fast-paced environment.
Site Reliability Engineer responsible for defining SLIs/SLOs, leading incident response, building observability (Prometheus/Grafana), automating toil, and driving production readiness for Runpod's AI cloud platform. Requires 5+ years SRE experience, strong Linux/distributed systems knowledge, and scripting skills.
Security Engineer responsible for endpoint security architecture, MDM administration (Jamf, Intune), compliance enforcement, and automation across macOS, Windows, iOS, and Android devices. Requires 3-6 years MDM experience, OSQuery/CrowdStrike expertise, and a security-first mindset.
Build and improve Cloudflare Workers Runtime: enhance performance, reliability, JavaScript and WebAssembly support for edge execution of customer code. Requires strong CS fundamentals, production ownership, and 2+ years C++/Rust experience.
Own the full product security lifecycle for space communications systems, from threat modeling and secure architecture to penetration testing, vulnerability management, cryptography, and compliance with FedRAMP, CMMC, and NIST standards. Requires 5+ years of product/application security leadership, deep expertise in SAST/DAST, secrets management, CI/CD hardening, and applied cryptography; TS/SCI clearance eligibility required.
Own vendor performance, SLAs, relationships, and governance for GPU hyperscalers and neocloud providers that power fal's generative media infrastructure. Requires 5+ years vendor management experience in AI/cloud/GPU infrastructure, with direct work on commercial relationships and performance accountability.
Lead and scale Runpod's core cloud and bare-metal infrastructure, including SRE, global/HPC networking, and distributed storage for massive GPU/AI workloads. Requires 8+ years operating large-scale distributed systems plus 7+ years leading infrastructure/SRE teams (managing managers).
Senior ML Engineer optimizing and productionizing LLMs and other models on Cloudflare's global serverless inference platform. Focus on inference performance, benchmarking, evaluation, and deployment at scale across heterogeneous GPUs and accelerators.