Latest Cloud Infrastructure jobs
Job results
Owns reliability, scalability, and observability of cloud infrastructure including compute, storage, and networking at massive scale. Drives SLOs, incident response, tooling, and mentors engineers; requires 15+ years experience with data centers and internet-scale operations.
Architects and builds high-performance Golang backend services for AI PaaS platform, focusing on distributed systems, GPU orchestration, model deployment, and cloud-native infrastructure. Requires 5+ years experience with expert Golang and Kubernetes proficiency.
Build end-to-end features for AI compute platform including React UIs, Python/Node.js APIs, data workflows, and compute integrations. Requires 5+ years full stack experience, React, Python/FastAPI or Node.js, and startup environment comfort.
Builds and manages platforms for AI model lifecycles, focusing on fine-tuning, training pipelines, and reinforcement learning for LLMs. Requires 8+ years in AI, advanced degree, and hands-on experience with generative AI techniques.
Builds and maintains platforms for fine-tuning, training, and managing LLMs including reinforcement learning pipelines and multi-node orchestration. Requires 4+ years in AI, hands-on LLM experience, and advanced degree in CS/Engineering.
Manages financial lifecycle of data center construction projects, including reviewing contractor pay apps, applying US GAAP for capital projects, contract reviews, month-end closes, and audit support. Requires Bachelor's in Accounting/Finance, CPA, 3+ years experience, and strong Excel proficiency.
Designs and implements SOX-compliant internal controls for core business processes like revenue, procurement, and financial close. Partners cross-functionally with Engineering, Finance, Sales, and Operations to ensure scalable, audit-ready systems in a high-growth AI infrastructure company. CPA required.
Owns architectural design, standards, and management for multi-site data center campuses, from concept through construction. Manages external teams, ensures code compliance (IBC/IFC/NFPA), and handles permitting for hyperscale facilities. Requires 7+ years experience, architecture degree, and proficiency in Revit/AutoCAD.
Leads technical strategy for power distribution, switchgear design, and protection systems in modular data centers. Oversees system analysis, code compliance, design reviews, and mentors engineers. Requires 10-15+ years experience, BSEE (MSEE/PE preferred), and deep power systems expertise.
Owns execution of employee lifecycle processes including onboarding, offboarding, promotions, and leaves in a specific region. Ensures reliable, compliant operations using tools like Rippling, with focus on employee experience and process improvements.
Owns civil engineering workstream for multi-site data center campuses, from site evaluation and design to permitting and construction. Requires 7+ years experience, PE license or pursuit, and proficiency in AutoCAD Civil 3D.
Owns project management for data center design and engineering programs, coordinating multidisciplinary teams, external consultants, and delivery partners. Builds templates, automation, and governance to drive execution across portfolio from design to construction handoff. Requires 7+ years experience and bachelor's in engineering/architecture/construction.
Build and maintain infrastructure for a large fleet of GPU servers, including provisioning, health monitoring, diagnostics, recovery, storage optimization, and Linux tuning for AI workloads. Requires 3+ years managing large-scale bare-metal/cloud fleets, strong Python and deep Linux expertise.
Builds and scales Rust-based services for edge-to-cloud distributed systems integrated with Kubernetes and major clouds (AWS, Azure, GCP). Requires 6+ years experience, preferably from FAANG/cloud providers, with strong distributed systems expertise.
Senior Data Scientist builds forecasting models, performs root cause analysis, analyzes cloud infrastructure usage, and champions data quality to support AI infrastructure operations and strategic decisions. Requires strong SQL/Python skills, analytical methods, and data visualization proficiency.
Builds and operates Upbound Spaces, a multi-control plane management platform using Go and Kubernetes. Troubleshoots production issues, develops features, contributes to open-source Crossplane, and ensures scalability and reliability in cloud environments.
Strategic Account Executive responsible for developing, managing, and closing enterprise sales for Upbound's AI-native infrastructure platform in North America. Partners cross-functionally to drive complex sales cycles with large G2K companies, focusing on cloud-native and AI solutions.
Senior sales professional driving revenue in Federal Civil government accounts through territory planning, pipeline management, contract negotiations, and building strategic relationships. Requires 5+ years selling technical solutions to US Federal customers with networking knowledge.
Leads Crusoe’s EMEA partnership strategy and execution, building startup, channel, and systems-integrator ecosystems to generate pipeline and bookings for its AI infrastructure business. Requires 7+ years of cloud or AI/ML partnerships experience, EMEA market expertise, and a bachelor’s degree.
Maintains general ledger integrity, owns full-cycle accounting, supports month-end/year-end closes, and ensures accurate financial reporting in a fast-scaling AI infrastructure company. Requires Bachelor's in Accounting/Finance, CPA, US GAAP expertise, and ERP proficiency.
Oversees general ledger, month-end/year-end closes, financial reporting, and internal controls while partnering cross-functionally and with auditors. Requires CPA, GAAP knowledge, ERP/Excel proficiency, and thrives in fast-paced scaling environments.
Sells Cloudflare solutions to high-growth digital native startups, focusing on new business acquisition, customer expansion, and renewals in NYC territory. Requires 5+ years B2B SaaS sales experience, technical aptitude, and proven quota attainment.
Maintains operational integrity of modular data centers through monitoring, mechanical/electrical maintenance, network troubleshooting, and critical repairs to minimize downtime for AI workloads. Requires hands-on expertise in IT systems, infrastructure like UPS/PDUs/HVAC, and safety protocols.
Designs and builds RF electronics hardware including amplifiers, filters, and support circuits for space communications ground networks. Collaborates across disciplines to optimize performance from concept to production; requires bachelor's in electrical engineering and 0-4+ years RF experience.
Builds and operates reliable, scalable AI infrastructure including observability, SLOs, incident response, automation, and performance tuning for ultra-low-latency serverless compute. Requires 3+ years SRE/DevOps experience with cloud, Kubernetes, programming (Go/Rust/Python), and observability tools.
Build and operate highly available distributed systems, microservices, and control-plane components for Kong’s managed gateway products. The role requires 5+ years of software experience, strong Golang expertise, cloud and Kubernetes experience, and deep knowledge of networking and API infrastructure.
Leads full-lifecycle engineering design of modular data centers, owning electrical, mechanical, thermal, and structural domains for AI infrastructure. Requires 10+ years in data center design, partner management expertise, and mastery of safety standards and CAD tools.
Executes physical deployment of data center whitespace infrastructure including racks, power systems, containment, and cabling. Oversees contractors, ensures compliance with designs and safety standards, and manages handoffs to operations. Requires 5+ years in data center construction and engineering degree.
Leads electrical design, optimization, and roadmap for prefabricated modular AI data centers (Crusoe Spark). Requires 5+ years in modular electrical systems, power distribution for AI compute, and cross-functional collaboration. In-office role in Denver with 10-20% travel.
Designs, reviews, and oversees installation/testing of mechanical systems like HVAC, piping, plumbing, and fire protection in data centers. Requires 3-5 years experience, bachelor's in Mechanical Engineering, and knowledge of building codes/standards.
Build and maintain infrastructure tooling for a large fleet of GPU servers, including provisioning, health monitoring, diagnostics, recovery, storage optimization, and Linux tuning to support AI workloads at scale. Requires 3+ years managing large server fleets, strong Python and deep Linux expertise.
Seasoned SRE owning reliability of Kubernetes-based production infrastructure at scale for a generative AI platform. Responsibilities include operating clusters, CI/CD, SLOs, monitoring, automation with AI, and driving improvements via chaos engineering. Requires 5+ years production experience with deep Kubernetes and observability expertise.
The Staff DevOps/SRE Engineer will define infrastructure strategy and SRE practices while building reliable, scalable systems for distributed, multi-cloud AI workloads. The role requires 8+ years of experience, deep Kubernetes and infrastructure-as-code expertise, and strong capabilities in automation, observability, and incident response.
Leads technical training programs for manufacturing teams, delivering hands-on instruction in fabrication, assembly, and electrical processes while managing program development, content creation, and employee certifications. Requires 5+ years manufacturing experience and training expertise.
Customer Success Manager responsible for building relationships, providing technical guidance on AI/ML and Kubernetes solutions, performance monitoring, training, and issue resolution to maximize customer value and retention. Requires bachelor's degree, proven customer success experience, and strong technical understanding of cloud, AI/ML.
Leads escalations for complex cloud incidents in AI infrastructure, designs reliability improvements for Kubernetes and GPU clusters, troubleshoots AI/ML workloads, and mentors engineers. Requires 8+ years in SRE/DevOps/HPC with deep Linux, networking, and customer expertise.
Leads internal audit function, overseeing audit lifecycle, risk assessments, and compliance across financial, operational, and IT controls. Provides strategic risk mitigation advice to executives and board, requiring 10+ years audit experience and strong leadership skills.
Administers and optimizes Atlassian Cloud tools like Jira and Confluence for enterprise collaboration, ITSM, and reporting. Customizes workflows, drives AI initiatives with Rovo, and ensures security/integrations. Requires 3+ years experience.
Leads HR strategies for the Real Estate team focused on data center development, driving organizational development, change management, employee growth, engagement, and retention. Partners with leadership on performance management, career development, and talent planning in a fast-paced environment.
Leads strategic sales for Cloudflare's major enterprise accounts, closing multi-million dollar platform deals. Requires 10+ years B2B sales experience, enterprise architecture expertise, C-suite engagement skills, and sales tools proficiency.
Technical expert driving presales for Cloudflare's Digital Native mid-market customers, leading demos, PoCs, and implementations while expanding accounts. Requires 6+ years customer-facing experience and deep internet technologies knowledge.
Senior Solutions Engineer drives technical sales for Cloudflare's majors accounts, building relationships with technical stakeholders, designing solutions for security and network needs, and evangelizing Cloudflare to executives.
Customer-facing technologist partnering with sales to deliver technical presentations, demos, and proofs-of-concept for Cloudflare solutions to enterprise customers. Requires experience in pre-sales technical roles with CDN, security, networking, or SaaS.
Strategic Technical Sourcer maps talent pools, builds proactive pipelines for engineering roles, and engages top passive technical candidates using data-driven strategies and specialized channels like GitHub and arXiv. Requires 2+ years sourcing experience in high-growth tech.
Strategic Talent Acquisition Partner owning full-cycle recruiting, workforce planning, and data-driven hiring strategy for a high-growth AI infrastructure company. Partners with leadership to anticipate talent needs and raise the hiring bar with 2+ years FLC experience required.
Leads multi-million-dollar sales cycles for AI infrastructure to foundation model companies and enterprises. Builds executive relationships, shapes GTM strategy, and drives revenue growth through complex deals and cross-functional collaboration.
Leads integration of electrical, thermal, mechanical, and networking systems in modular data centers for high-density AI compute. Ensures compatibility with power sources, optimizes thermal management, and supports transition to liquid cooling with 6+ years systems engineering experience.
Build and own full-stack projects to enhance Render's cloud platform, focusing on user-facing features for effortless deployments. Requires 6+ years experience across the stack with web frameworks, APIs, databases, and multiple languages.
Owns end-to-end product design for web and native mobile features, maintains design systems, prototypes rapidly with Figma and AI tools, conducts lightweight user research, and collaborates with PM/Eng in a high-velocity startup. Requires strong UI craft, shipped portfolio, and cross-platform experience.
Builds end-to-end features for an AI compute infrastructure platform, spanning React interfaces, Python or Node.js APIs, data models, integrations, and compute workflows. The role requires 5+ years of full-stack experience and strong ownership across frontend and backend systems.