Skip to content

Latest DevOps / SRE jobs at Crusoe

Search
Location
38 jobs

Job results

Crusoe

Senior Production Engineer, Compute

CrusoeSunnyvale, CA

Senior Production Engineer responsible for optimizing Crusoe's virtualization, hypervisor, and Linux kernel stack to deliver high-performance AI and HPC compute infrastructure. Requires 5+ years experience with kernel internals, KVM/QEMU, low-level debugging, and performance tuning for GPUs and DPUs.

170k – 205k/yrOn-site5+ YOEDevOps / SRE
Crusoe

Senior Production Engineer, SDN

CrusoeSunnyvale, CA

Senior Production Engineer focused on building automation, self-healing tools, and reliability for Crusoe's SDN infrastructure that powers AI and HPC workloads. Requires 5+ years experience automating network provisioning with deep expertise in SDN platforms, Linux networking, Kubernetes CNIs, and protocols like BGP.

170k – 205k/yrOn-site5+ YOEDevOps / SRE
Crusoe

Senior Production Engineer, Storage

CrusoeSunnyvale, CA

Build and optimize distributed, fault-tolerant cloud storage systems (block, file, object) for Crusoe's AI/HPC infrastructure. Ensure high availability, performance, and reliability through automation, incident response, and collaboration with hardware/kernel teams. Requires 5+ years in storage engineering/SRE with deep Linux and IaC expertise.

170k – 205k/yrOn-site5+ YOEDevOps / SRE
Crusoe

Senior Production Engineer, Managed Cloud

CrusoeSan Francisco, CA

Senior Production Engineer responsible for designing, operating, and optimizing reliable managed AI cloud services focused on scaling LLM workloads, defining SLIs/SLOs, building observability, and resolving issues in distributed systems for Crusoe's AI infrastructure.

170k – 205k/yrOn-site5+ YOEDevOps / SRE
Crusoe

Senior Software Engineer, Developer Experience

CrusoeSan Francisco, CA +1

Senior Software Engineer on the Developer Experience team building internal tools, libraries, CI/CD pipelines, and paved paths to accelerate engineering productivity and eliminate toil across the full SDLC at Crusoe. Requires strong Go, Kubernetes, DevOps/SRE background and experience creating developer infrastructure.

172k – 209k/yrOn-site5+ YOEDevOps / SRE
Crusoe

Principal Engineer, CAPE

CrusoeSan Francisco, CA

Principal Engineer building Crusoe's self-driving Conductor platform for AI infrastructure. Own closed-loop autonomy, predictive failure detection, unified observability, energy-aware scheduling, and goodput optimization across tens of thousands of GPUs. Requires 10+ years in large-scale distributed systems, HPC/GPU infrastructure, and observability.

285k – 335k/yrOn-site10+ YOEDevOps / SRE
Crusoe

Senior Performance Engineer

CrusoeSan Francisco, CA

Senior Performance Engineer responsible for Linux kernel optimization, system benchmarking, and low-level performance tuning to enhance Crusoe's AI cloud infrastructure. Requires deep Linux kernel expertise, proficiency in Go/C/C++, and hands-on experience with performance optimization in complex environments.

170k – 205k/yrOn-site5+ YOEDevOps / SRE
Crusoe

Staff Infrastructure Engineer

CrusoeSan Francisco, CA +1

Staff Cloud Infrastructure Engineer responsible for managing Crusoe’s cloud fleet operations, building IaC automation for bare-metal server provisioning/reprovisioning, scaling deployments, GPU hardware troubleshooting, and transitioning to Kubernetes/containerized workflows.

208k – 253k/yrOn-site7+ YOEDevOps / SRE
Crusoe

Staff Production Engineer, Compute

CrusoeSan Francisco, CA

Staff Production Engineer responsible for developing automation/observability, scaling virtualization (KVM/QEMU), optimizing Linux kernel performance, and supporting AI/HPC workloads on CPU/GPU/DPU hardware. Requires 8+ years in Linux systems engineering, kernel internals, and virtualization.

209k – 253k/yrOn-site8+ YOEDevOps / SRE
Crusoe

Staff Software Engineer, Cloud Monitoring Service

CrusoeSan Francisco, CA

Lead the design and evolution of Crusoe Cloud's large-scale telemetry and observability systems for metrics and logs. Own high-throughput distributed pipelines from edge collection through ingestion, storage, and low-latency querying while ensuring scalability, reliability, and multi-tenancy.

215k – 260k/yrOn-site8+ YOEDevOps / SRE
Crusoe

Senior Microsoft Cloud Infrastructure Engineer

CrusoeSan Francisco, CA

Senior Cloud Infrastructure Engineer owning design, implementation, and management of Microsoft 365, Entra ID, Azure, and Azure Arc hybrid infrastructure. Requires 8+ years infrastructure experience with deep Azure/M365 expertise, IaC, Windows Server admin, and hands-on data center hardware work.

160k – 195k/yrOn-site8+ YOEDevOps / SRE
Crusoe

Senior Staff Engineer, Platform R&D

CrusoeSan Francisco, CA

Senior individual contributor embedded in Crusoe's Managed Platform Services team to accelerate delivery through rapid AI-augmented R&D, prototyping, and cross-domain technical leadership. Requires 10+ years experience with systems languages and cloud-native infrastructure.

245k – 295k/yrOn-site10+ YOEDevOps / SRE
Crusoe

Staff Software Engineer, Developer Experience

CrusoeSan Francisco, CA +1

Staff-level engineer building developer tools, infrastructure, and automation to accelerate Crusoe engineering productivity. Requires Go, Kubernetes, CI/CD, and strong DevOps/SRE experience.

209k – 253k/yrOn-site7+ YOEDevOps / SRE
Crusoe

Senior Staff Network Engineer, Operations

CrusoeSan Francisco, CA

The Senior Staff Network Operations Engineer will own production reliability for Crusoe's global network, including edge, backbone, data center fabric, and GPU cluster interconnects. This role involves leading incident response, driving root cause analysis, defining SLIs/SLOs, and setting operational standards to maintain hyperscale AI infrastructure health.

225k – 275k/yrOn-site12+ YOEDevOps / SRE
Crusoe

Senior Staff Network Engineer, Automation

CrusoeSan Francisco, CA

Senior technical leader owning Crusoe's network automation platform, source of truth, intent-based config systems, and self-healing workflows across hyperscale multi-vendor fabrics. Requires 12+ years of production network automation experience with deep expertise in Python/Go, model-driven telemetry, and observability at 10K+ device scale.

245k – 295k/yrOn-site12+ YOEDevOps / SRE
Crusoe

Senior Production Engineer

CrusoeSan Francisco, CA +1

As a Senior Production Engineer, you will ensure the reliability and scalability of Crusoe’s AI-optimized cloud platform, focusing on designing and operating managed AI services for LLM workloads. You will build automation and reliability tooling, define SLIs/SLOs, and optimize large-scale training and inference clusters.

209k – 253k/yrOn-siteDevOps / SRE
Crusoe

Senior Staff Network Engineer, Deployment

CrusoeSan Francisco, CA +2

Senior technical leader owning global network infrastructure deployment strategy, automation platforms, and standards for Crusoe's hyperscale AI data centers. Requires 12+ years of large-scale data center deployment experience with deep expertise in Arista, Juniper, and NVIDIA platforms.

225k – 275k/yrOn-site12+ YOEDevOps / SRE
Crusoe

Staff Software Engineer, Managed Orchestration (Managed Kubernetes)

CrusoeSan Francisco, CA +1

Staff Software Engineer designs, builds, and scales managed Kubernetes and AI training clusters, focusing on reliability, performance, and orchestration using Go, Terraform, and GCP. Oversees architecture, CI/CD pipelines, and critical infrastructure projects requiring 8+ years experience.

220k – 250k/yrOn-site8+ YOEDevOps / SRE
Crusoe

Senior Production Engineer, Operational Excellence

CrusoeSan Francisco, CA +1

Senior Production Engineer ensures reliability, scalability, and performance of GPU cloud infrastructure powering AI workloads. Drives observability, incident response, automation, and operational improvements in large-scale distributed systems.

172k – 209k/yrOn-site5+ YOEDevOps / SRE
Crusoe

Senior Virtualization Validation Engineer

CrusoeSan Francisco, CA +1

Validates large-scale multi-node GPU clusters using QEMU and Cloud Hypervisor, focusing on interconnects like NVLink/InfiniBand, collective communications (NCCL/RCCL), and performance in virtualized AI/HPC environments. Requires 5+ years experience, virtualization expertise, and Linux kernel knowledge.

173k – 210k/yrOn-site5+ YOEDevOps / SRE
Crusoe

Staff Instrumentation & Controls Engineer, Deployment

CrusoeChildress, TX +1

Leads deployment, integration, and startup of BMS/EPMS/SCADA systems in data centers, ensuring seamless operation of HVAC, electrical, and monitoring infrastructure. Oversees contractors, troubleshoots protocols like BACnet/Modbus, and validates systems via FAT/SAT. Requires Bachelor's in engineering and hands-on data center automation experience.

148k – 170k/yrOn-siteDevOps / SRE
Crusoe

Electrical Field Engineer - Data Center

CrusoeTexas +6

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.

196k – 235k/yrRemote5+ YOEDevOps / SRE
Crusoe

Senior API Integration Engineer

CrusoeSunnyvale, CA +1

Leads design and delivery of enterprise API integrations using Workato for People Tech ecosystem, automating workflows across ERP, CRM, HCM systems. Requires 7+ years experience with Workato recipes, iPaaS patterns, API security, and stakeholder collaboration.

165k – 200k/yrOn-site7+ YOEDevOps / SRE
Crusoe

Senior Staff Software Engineer, Managed Orchestration

CrusoeSan Francisco, CA

Leads architecture and development of scalable managed Kubernetes and AI orchestration systems, providing technical direction for cloud infrastructure reliability and performance. Requires 10+ years in software engineering with deep expertise in Go, Kubernetes, and large-scale systems.

238k – 288k/yrOn-site10+ YOEDevOps / SRE
Crusoe

Staff Software Engineer, Systems Engineering Focus

CrusoeSan Francisco, CA

Designs, builds, and scales customer-facing managed services with a focus on edge agents running on customer infrastructure. Provides technical oversight for high-reliability systems using eBPF, Kubernetes, and low-level Linux metrics; leads cross-team collaboration and mentors engineers.

210k – 255k/yrOn-siteDevOps / SRE
Crusoe

Senior Staff Engineer, Cloud Site Operations

CrusoeSan Francisco, CA +1

Leads technical architecture for data center operations, overseeing global ticket queues, fleet supportability, power topology, resilience planning, and hardware failure escalations for AI infrastructure. Requires 10+ years in data center ops or HPC with deep NVIDIA GPU expertise.

179k – 218k/yrOn-site10+ YOEDevOps / SRE
Crusoe

Data Center Systems Engineer, R&D

CrusoeDenver, CO

Leads R&D for next-generation data center architectures, evaluating emerging technologies across power, cooling, compute, and facilities. Defines reference designs, guides pilot-to-hyperscale progression, and influences cross-functional teams. Requires 15+ years experience and systems thinking across engineering domains.

Salary not listedOn-site15+ YOEDevOps / SRE
Crusoe

Senior Staff Storage Systems Administrator

CrusoeSan Francisco, CA

Leads architecture, operation, and vendor strategy for petabyte-scale storage systems optimized for AI/HPC workloads in sustainable cloud infrastructure. Requires 10+ years experience with enterprise storage, scripting, and RFP/vendor management.

170k – 215k/yrOn-site10+ YOEDevOps / SRE
Crusoe

Network Architect

CrusoeSan Francisco, CA

Defines and governs network architecture, security strategy, and standards for data centers, power plants, and corporate environments. Requires 8+ years network engineering with expertise in IT, cloud, and industrial systems like IoT/OT.

195k – 225k/yrOn-site8+ YOEDevOps / SRE
Crusoe

Principal Production Engineer

CrusoeSan Francisco, CA +1

Owns reliability, scalability, and observability of cloud infrastructure including compute, storage, and networking at massive scale. Drives SLOs, incident response, tooling, and mentors engineers; requires 15+ years experience with data centers and internet-scale operations.

261k – 326k/yrOn-site15+ YOEDevOps / SRE
Crusoe

Infrastructure Deployment Engineer

CrusoeSpringfield, OH

Executes physical deployment of data center whitespace infrastructure including racks, power systems, containment, and cabling. Oversees contractors, ensures compliance with designs and safety standards, and manages handoffs to operations. Requires 5+ years in data center construction and engineering degree.

120k – 150k/yrOn-site5+ YOEDevOps / SRE
Crusoe

Staff Modular Data Center Engineer

CrusoeDenver, CO

Leads electrical design, optimization, and roadmap for prefabricated modular AI data centers (Crusoe Spark). Requires 5+ years in modular electrical systems, power distribution for AI compute, and cross-functional collaboration. In-office role in Denver with 10-20% travel.

168k – 192k/yrOn-site5+ YOEDevOps / SRE
Crusoe

Principal Systems Software Engineer

CrusoeSan Francisco, CA +1

Leads architecture of next-generation AI infrastructure, unifying BMaaS, IaaS, and CaaS with focus on high-performance I/O paths, kernel optimizations, and GPU workloads. Requires 12+ years hyperscale experience, deep Linux/virtualization expertise, and hardware-software co-design skills.

260k – 340k/yrOn-site12+ YOEDevOps / SRE
Crusoe

Atlassian Cloud Engineer

CrusoeSan Francisco, CA

Administers and optimizes Atlassian Cloud tools like Jira and Confluence for enterprise collaboration, ITSM, and reporting. Customizes workflows, drives AI initiatives with Rovo, and ensures security/integrations. Requires 3+ years experience.

115k – 135k/yrOn-site3+ YOEDevOps / SRE
Crusoe

Data Center Systems Engineer (MDC)

CrusoeDenver, CO +1

Leads integration of electrical, thermal, mechanical, and networking systems in modular data centers for high-density AI compute. Ensures compatibility with power sources, optimizes thermal management, and supports transition to liquid cooling with 6+ years systems engineering experience.

149k – 170k/yrOn-site6+ YOEDevOps / SRE
Crusoe

Staff Production Engineer (Operational Excellence)

CrusoeSan Francisco, CA +1

Leads reliability and operational excellence for Crusoe's GPU cloud platform, driving incident response, observability with Prometheus/Grafana, automation to reduce toil, and SLOs for AI/HPC workloads. Requires 8+ years in SRE/production engineering, Kubernetes, Linux, and programming in Python/Go.

209k – 253k/yrOn-site8+ YOEDevOps / SRE
Crusoe

Staff Network Engineer, Deployment

CrusoeSan Francisco, CA +3

Leads physical and logical deployment of global network infrastructure for AI data centers, including rack/stack, cabling, automation with Python/Ansible, testing, and partner coordination. Requires 8+ years experience with Arista, Juniper, Mellanox, BGP/EVPN, and physical layer expertise.

193k – 234k/yrOn-site8+ YOEDevOps / SRE
Crusoe

Senior Software Engineer, Managed Orchestration (Managed Kubernetes)

CrusoeSan Francisco, CA +1

Build and scale managed Kubernetes and AI training clusters, developing operators, controllers, and infrastructure using Go, Terraform, and GCP. Design reliable, high-performance systems competing with GKE/EKS, with 5+ years experience required.

180k – 210k/yrOn-site5+ YOEDevOps / SRE