Skip to content
CrusoeCrusoeSan Francisco, CA

Senior Software Engineer

Develop diagnostics, observability, automation, and repair tooling for Crusoe’s large-scale GPU infrastructure and data centers. The role requires 4–6 years of software engineering experience, distributed systems expertise, and proficiency in Go, Python, Java, or Rust.

$170k – $205k/yr
On-site5+ YOEFullstack Engineering

About the job

Responsibilities

  • Develop deep-level diagnostics and troubleshooting for hardware faults within GPU racks and high-density compute systems.
  • Build troubleshooting and automation tooling for NVIDIA A100, H200, GB200, B200, and AMD 350X/355X GPU platforms.
  • Develop automation and AI agents for component-level diagnosis and remediation of failed or degraded hardware.
  • Partner with data center operations to create tooling and AI agents for managing critical environments.
  • Develop post-repair validation and testing tools, including burn-in, PyTorch, and NVIDIA NCCL, to ensure system stability and performance.
  • Own deployment, monitoring, and operational support for developed tooling to maximize GPU fleet availability and performance.
  • Develop automation and operational tooling for facility power management and direct liquid-cooling hardware systems.

Requirements

  • 4–6 years of software engineering experience.
  • Ability to identify problems, rapidly develop scalable solutions, and ship them.
  • Ability to contribute to critical or complex technical initiatives.
  • Ability to set technical direction for a project and execute it.
  • Expertise in distributed systems, reliability, and cloud platforms.
  • Strength in at least one programming language: Go, Python, Java, or Rust.
  • Strong analytical and problem-solving skills.
  • Excellent communication and collaboration skills.
  • Ability to work independently and as part of a team.

Nice to Have

  • Experience with Temporal and Kubernetes.
  • Experience working directly with hardware vendors.
  • Experience with large-scale GPU fleet operations or hyperscale data center environments.

Compensation and Benefits

  • Compensation range of $170,000–$205,000 plus bonus.
  • Restricted Stock Units included in all offers.
  • Health insurance options including HDHP and PPO, vision, and dental coverage.
  • Employer HSA contributions.
  • Paid parental leave.
  • Company-paid life insurance and short- and long-term disability coverage.
  • Teladoc.
  • 401(k) with a 100% match up to 4% of salary.
  • Generous paid time off and holiday schedule.
  • Cell phone reimbursement.
  • Tuition reimbursement.
  • Calm app subscription.
  • MetLife Legal.
  • Company-paid commuter benefit of $50 per pay period.

Skills

GoPythonJavaRustKubernetesTemporalGCPInfrastructure As CodeDistributed SystemsPyTorchNvidia NcclGpu Fleet OperationsDiagnostics AutomationLiquid Cooling