Skip to content
CrusoeCrusoe

Staff Software Engineer

Build diagnostics, automation, observability, and repair tooling for Crusoe’s large-scale GPU fleet and data centers. The role requires software engineering expertise in distributed systems, reliability, cloud platforms, and at least one of Go, Python, Java, or Rust.

About the job

Responsibilities

  • Develop deep-level diagnostics and troubleshooting for hardware faults within GPU racks and high-density compute systems.
  • Build troubleshooting and automation tooling for NVIDIA A100, H200, GB200, B200, and AMD 350X/355X GPU platforms.
  • Develop automation and AI agents for component-level diagnosis and remediation of failed or degraded hardware.
  • Partner with data center operations to build tooling and AI agents for managing critical environments.
  • Develop post-repair validation and testing tools, including burn-in, PyTorch, and NVIDIA NCCL, to ensure system stability and performance.
  • Own deployment, monitoring, and operational support for developed tooling to maximize GPU fleet availability and performance.
  • Develop automation and operational tooling for facilities management, power, and direct liquid-cooling hardware systems.
  • Identify problems, rapidly develop scalable solutions, and ship them.
  • Set technical direction for specific projects and execute independently or collaboratively.

Requirements

  • Software engineering experience.
  • Expertise in distributed systems, reliability, and cloud platforms.
  • Experience with Kubernetes, infrastructure as code, and Google Cloud.
  • Proficiency in at least one programming language: Go, Python, Java, or Rust.
  • Strong analytical and problem-solving skills.
  • Excellent communication and collaboration skills.
  • Ability to work independently and assist with critical or complex technical initiatives.

Nice-to-Haves

  • Experience with Temporal and Kubernetes.
  • Experience working directly with hardware vendors.
  • Experience operating large-scale GPU fleets or hyperscale data center environments.

Compensation and Benefits

  • Compensation range of $215,000–$260,000 plus bonus.
  • Restricted Stock Units included in all offers.
  • Health insurance options including HDHP and PPO, vision, and dental coverage.
  • Employer HSA contributions.
  • Paid parental leave.
  • Paid life insurance and short- and long-term disability coverage.
  • Teladoc.
  • 401(k) with a 100% match up to 4% of salary.
  • Paid time off and holidays.
  • Cell phone and tuition reimbursement.
  • Calm app subscription.
  • MetLife Legal.
  • Company-paid commuter benefit of $300 per month.

Skills

Gpu Infrastructure, Distributed Systems, Reliability Engineering, Kubernetes, Infrastructure As Code, GCP, Go, Python, Java, Rust, Temporal, PyTorch, Nvidia Nccl, AI Agents, Direct Liquid Cooling

Reddit

Reddit

San Francisco, CA

Staff Site Reliability Engineer - Site Experience
$217k+/yrOn-site8+ YOEDevOps / SRE

Leads reliability engineering for Reddit’s critical user-facing systems, improving availability, scalability, performance, automation, and incident response at internet scale. Requires 8+ years operating distributed systems and strong expertise in programming, observability, high availability, and production troubleshooting.

Reddit

Reddit

San Francisco, CA

Staff Site Reliability Engineer, Ads
$217k+/yrRemote8+ YOEDevOps / SRE

Provides technical leadership for reliability, scalability, and operational excellence across Reddit’s advertising systems. The role requires 8+ years operating large-scale distributed systems, strong software engineering skills, and expertise in cloud-native architectures, observability, and incident response.

Reddit

Reddit

United States

Staff Software Engineer, Observability
$217k+/yrRemote7+ YOEDevOps / SRE

Build and operate Reddit’s internet-scale observability platform across monitoring, logging, and distributed tracing. The role requires 7+ years of infrastructure or software engineering experience, distributed systems expertise, and strong Kubernetes and troubleshooting skills.

Airbnb

Airbnb

United States

Staff Software Engineer, Service Tools
$212k+/yrRemote9+ YOEDevOps / SRE

Leads technical direction for Airbnb’s service developer tooling platform, spanning AI-assisted development, JVM build infrastructure, testing, modernization, and observability. Requires 9+ years of industry experience, strong backend and distributed-systems expertise, and the ability to influence organizations and deliver multi-quarter infrastructure initiatives.

Temporal

Temporal

United States

Staff Software Engineer, Traffic
$212k+/yrRemote8+ YOEDevOps / SRE

Leads the design and development of scalable, secure network traffic systems and cloud infrastructure. The role requires 8+ years of coding experience, strong distributed-systems and concurrency expertise, and deep knowledge of networking and performance optimization.