Skip to content
CrusoeCrusoe

Staff Production Engineer, Compute

Staff Production Engineer responsible for developing automation/observability, scaling virtualization (KVM/QEMU), optimizing Linux kernel performance, and supporting AI/HPC workloads on CPU/GPU/DPU hardware. Requires 8+ years in Linux systems engineering, kernel internals, and virtualization.

About the job

What You'll Be Working On

  • Develop automation and observability tools to monitor Crusoe’s compute infrastructure, spanning from the kernel to orchestration layers.
  • Support and scale the company’s virtualization stack, including technologies such as KVM, QEMU, and other hypervisors.
  • Collaborate with Linux kernel and hardware teams to identify and resolve performance bottlenecks, driver issues, and optimize hardware offloads.
  • Optimize performance for AI and HPC workloads across CPU, GPU, and DPU/NIC resources.
  • Participate in root cause analysis for kernel crashes, hardware-software integration problems, and performance regressions.
  • Integrate hypervisor-level enhancements to improve guest VM reliability and workload isolation.
  • Tune kernel subsystems such as the process scheduler, NUMA configuration, memory management, and interrupt handling.
  • Work closely with platform teams to implement and validate support for emerging compute hardware, including SmartNICs, BlueField devices, and TPUs.

What You’ll Bring to the Team

  • 8+ years of professional experience in Compute SRE, Linux system engineering, or compute infrastructure roles.
  • Strong proficiency in Linux kernel internals, with exposure to scheduler, memory allocation, and driver subsystems.
  • Experience with virtualization architectures and technologies such as KVM, Xen, QEMU, or VMware.
  • Familiarity with SmartNICs/DPUs (e.g., NVIDIA CX6/7, BlueField-3) and kernel bypass techniques.
  • Expert-level skills in at least one programming language: Go, C or Rust.
  • Experience with system-level debugging, including kdump, kexec, and kernel panic analysis.
  • Proficiency in Infrastructure as Code tooling and CI/CD practices for bare-metal or cloud infrastructure.
  • Strong understanding of compute scheduling, resource management, and high-throughput networking.

Bonus Points

  • Experience porting or maintaining custom Linux distributions or kernels for specific platforms.
  • Exposure to AI model infrastructure and workload orchestration across GPU clusters.
  • Contributions to Linux kernel, KVM, or other low-level infrastructure projects.

Benefits

  • Competitive compensation
  • Restricted Stock Units
  • Paid time off & paid holidays
  • Comprehensive health, dental & vision insurance
  • Employer contributions to HSA account
  • Paid parental leave
  • Paid life insurance, short-term and long-term disability
  • Professional development & tuition reimbursement
  • Mental health & wellness support
  • Commuter benefits (parking & transit)
  • Cell phone stipend
  • 401(k) Retirement plan with company match up to 4% of salary
  • Volunteer time off

Skills

Linux Kernel, Kvm, Qemu, Xen, VMware, Smartnics, Dpus, Bluefield, Go, C, Rust, Kdump, Kexec, Infrastructure As Code, CI/CD

Airbnb

Airbnb

United States

Staff Software Engineer, Service Tools
$212k+/yrRemote9+ YOEDevOps / SRE

Leads technical direction for Airbnb’s service developer tooling platform, spanning AI-assisted development, JVM build infrastructure, testing, modernization, and observability. Requires 9+ years of industry experience, strong backend and distributed-systems expertise, and the ability to influence organizations and deliver multi-quarter infrastructure initiatives.

Temporal

Temporal

United States

Staff Software Engineer, Traffic
$212k+/yrRemote8+ YOEDevOps / SRE

Leads the design and development of scalable, secure network traffic systems and cloud infrastructure. The role requires 8+ years of coding experience, strong distributed-systems and concurrency expertise, and deep knowledge of networking and performance optimization.

Crusoe

Crusoe

San Francisco, CA
Staff Software Engineer
$215k+/yrOn-site7+ YOEDevOps / SRE

Build diagnostics, automation, observability, and repair tooling for Crusoe’s large-scale GPU fleet and data centers. The role requires software engineering expertise in distributed systems, reliability, cloud platforms, and at least one of Go, Python, Java, or Rust.

Reddit

Reddit

San Francisco, CA

Staff Site Reliability Engineer - Site Experience
$217k+/yrOn-site8+ YOEDevOps / SRE

Leads reliability engineering for Reddit’s critical user-facing systems, improving availability, scalability, performance, automation, and incident response at internet scale. Requires 8+ years operating distributed systems and strong expertise in programming, observability, high availability, and production troubleshooting.

Reddit

Reddit

San Francisco, CA

Staff Site Reliability Engineer, Ads
$217k+/yrRemote8+ YOEDevOps / SRE

Provides technical leadership for reliability, scalability, and operational excellence across Reddit’s advertising systems. The role requires 8+ years operating large-scale distributed systems, strong software engineering skills, and expertise in cloud-native architectures, observability, and incident response.