Latest Cloud Infrastructure jobs
Job results
Build and operate scalable infrastructure powering FlexAI’s AI and PaaS platform. The role focuses on Kubernetes, infrastructure as code, CI/CD, observability, incident response, and reliability practices, requiring 4+ years of DevOps, SRE, or infrastructure engineering experience.
Build and lead high-performance Golang backend services for an AI compute and PaaS platform. The role requires 8+ years of backend or infrastructure experience, distributed-systems expertise, and strong cloud-native, Kubernetes, and container orchestration skills.
Build and operate TypeScript backend microservices for Kong’s billing platform, covering subscriptions, entitlements, usage metering, provisioning, and payment integrations. The role requires 3+ years of backend engineering experience and expertise in TypeScript, Node.js, relational databases, distributed systems, and Kubernetes.
Drives enterprise sales of cloud and end-user computing solutions, partnering with Microsoft, channel partners, and customers to develop opportunities and close deals. Requires 5–7 years of technology sales experience, strong VDI/EUC knowledge, and proven quota attainment.
Develops test automation software and hardware fixtures for electronics and RF systems in satellite ground infrastructure. Requires 0-4 years experience, bachelor's in electrical engineering, strong programming in Python/C++/Rust, and hands-on test equipment skills.
Leads end-to-end delivery of data center infrastructure and AI cluster deployments, coordinating construction, networking, hardware, and operations teams across sites. Requires 5+ years in data center programs, AI/GPU expertise, and cross-functional influence in ambiguous environments.
Leads the development and evolution of the Application Security program, integrating security practices into SDLC through threat modeling, code reviews, and pen testing. Collaborates with engineering teams on web, mobile, and API security for a SaaS platform; requires 10+ years experience building AppSec from inception.
Hands-on network engineer deploying and validating large-scale AI datacenter fabrics, configuring switches, troubleshooting physical/optical layers, and coordinating cross-functional teams. Requires 3-7 years datacenter experience and 70-80% travel to onsite locations.
Designs and builds secure frameworks for AI workflows, microservices, and developer tools. Conducts security reviews and implements auth systems for SaaS offerings, requiring 7+ years experience in Golang/Node.js and distributed systems security.
Leads strategy and execution of large-scale OSP/ISP network infrastructure for AI data centers, managing full project lifecycles, vendors, budgets, and compliance. Requires 10+ years telecom construction leadership, bachelor's in engineering/construction, and expertise in fiber systems and permitting.
Senior Customer Success Manager responsible for building relationships, providing technical guidance on cloud AI/ML and Kubernetes solutions, driving adoption, and ensuring customer satisfaction and ROI for Crusoe's AI infrastructure offerings. Requires bachelor's in CS/Engineering, technical proficiency in AI/ML/cloud, and proven customer success experience.
Lead product vision and roadmap for Render's infrastructure platform supporting millions of developers. Requires 8+ years in product management focused on developer tools, infrastructure, or data products, with strong AI and developer experience interest.
Designs, implements, and optimizes global distributed systems for space communications, including control planes, APIs, edge hardware control, high-bandwidth networking, monitoring, and simulations. Requires 5+ years experience in software development with focus on cloud, embedded, and networking technologies.
Nurture and grow relationships with existing Runpod customers, focusing on retention, renewals, upsells of AI infrastructure, and driving net revenue retention through account planning, business reviews, and proactive health monitoring. Requires 2-5 years in post-sales roles with strong technical and relationship skills.
Founding Software Engineer designs and builds distributed systems for software deployment orchestration across complex enterprise environments. Requires 3+ years experience, full-stack fluency across languages and tools, and strong ownership to shape technical architecture with founders.
Provides advanced technical support for GPU-powered cloud infrastructure, diagnosing VM, hardware, storage, networking, and scaling issues while coordinating with global engineering teams. Requires at least five years of technical support experience, strong Linux and CLI skills, and familiarity with Kubernetes, cloud platforms, and HPC technologies.
Leads physical and logical deployment of global network infrastructure for AI data centers, including rack/stack, cabling, automation with Python/Ansible, testing, and partner coordination. Requires 8+ years experience with Arista, Juniper, Mellanox, BGP/EVPN, and physical layer expertise.
Builds and scales cloud infrastructure for Render's developer platform, focusing on container orchestration, networking, storage, and AI workloads. Requires 5+ years experience with Kubernetes, IaC tools like Terraform/Pulumi/Ansible, and production systems at scale.
Designs, develops, and optimizes Bluetooth Low Energy transport layer for peer-to-peer data sync in distributed systems. Requires 7+ years experience with deep BLE expertise and mobile development in C++/C/Kotlin for iOS/Android, plus Rust willingness.
Builds internal security tooling, implements detection systems, assesses vulnerabilities, and partners with engineering teams to enhance Render's security posture. Requires 6+ years in software engineering or security with experience in secure web apps and vulnerability analysis.
Designs, builds, and maintains usage-based billing systems, integrating third-party platforms like Orb, Metronome, and Stripe. Requires 3+ years experience with Go, databases, and cross-functional collaboration for accurate, scalable billing at Render.
Designs, builds, and optimizes distributed cloud storage systems for AI/HPC workloads, focusing on high-performance filesystems, block/object storage, and Linux subsystems. Requires deep expertise in scalable storage infrastructure and languages like Go, C++, Rust.
Designs and evolves production Ceph storage clusters, builds APIs and orchestration services for block/object storage using Go and gRPC. Requires experience with distributed systems, filesystems like ZFS/BTRFS, and building scalable infrastructure.
Designs and optimizes storage backends, transaction management, and CRDT operations for Ditto's embedded edge database on resource-constrained devices. Requires 5+ years experience with database internals, Rust or C/C++, and storage engines like SQLite or RocksDB.
Build and scale fal's core Python/Rust distributed platform for AI workload orchestration, scheduling, GPU autoscaling, and low-latency global inference. Requires 3+ years building production distributed systems with deep knowledge of consensus, fault tolerance, and observability.
Build and deploy interactive user experiences for fal's generative media platform, owning full-stack features from concept to launch with a focus on model playgrounds. Requires 5+ years full-stack experience with TypeScript, Python, Postgres, and Next.js.
Build end-to-end full stack products for Render's developer cloud platform, owning features from design to deployment. Requires 6+ years experience across the stack with web frameworks, APIs, databases, and distributed systems.
Build and scale managed Kubernetes and AI training clusters, developing operators, controllers, and infrastructure using Go, Terraform, and GCP. Design reliable, high-performance systems competing with GKE/EKS, with 5+ years experience required.
Leads advanced penetration testing, red team operations, and AI/ML security research to secure applications, infrastructure, and distributed AI systems including LLMs and Kubernetes environments. Requires 8-10 years offensive security experience and strong software engineering in Go/Python/Rust.
Develops and maintains ServiceNow applications, integrates with cloud platforms and tools like AWS, GCP, Azure, handles custom scripting, business rules, and customer implementations. Requires 5+ years in JS/ServiceNow development and ITOM experience.
Builds and operates global datacenters bridging physical infrastructure to software, focusing on performance, resilience, and reliability. Requires experience with Golang, gRPC, Postgres, OS primitives, oncall rotations, and delivering 0→N projects remotely.
Builds developer adoption through content, community, and open source contributions while gathering market feedback to shape product strategy. Requires strong communication, technical versatility across stacks, and comfort in startup ambiguity.
Senior Software Engineer building and scaling cloud infrastructure for AI/ML workloads on distributed systems across global data centers. Requires 3+ years experience with strong Go skills; distributed systems, containerization, and ML deployment experience preferred.
Builds low-level infrastructure software from first principles to power a cloud platform, focusing on OS primitives, distributed systems, and scalable gRPC services in Golang/Rust for high developer leverage.
Build end-to-end product features from dashboard UI to backend workflows using Temporal and microservices. Develop TypeScript + GraphQL APIs and contribute to Rust open-source tools for deployment infrastructure.