# Staff Software Engineer

**Company:** [Crusoe](https://hotfix.jobs/companies/crusoe)
**Location:** San Francisco, CA, Sunnyvale, CA
**Role:** DevOps / SRE
**Salary:** $215k – $260k/yr
**Experience:** 7+ years
**Skills:** Gpu Infrastructure, Distributed Systems, Reliability Engineering, Kubernetes, Infrastructure As Code, GCP, Go, Python, Java, Rust, Temporal, PyTorch, Nvidia Nccl, AI Agents, Direct Liquid Cooling
**Posted:** 2026-08-12

> Build diagnostics, automation, observability, and repair tooling for Crusoe’s large-scale GPU fleet and data centers. The role requires software engineering expertise in distributed systems, reliability, cloud platforms, and at least one of Go, Python, Java, or Rust.

## Job Description

## Responsibilities
- Develop deep-level diagnostics and troubleshooting for hardware faults within GPU racks and high-density compute systems.
- Build troubleshooting and automation tooling for NVIDIA A100, H200, GB200, B200, and AMD 350X/355X GPU platforms.
- Develop automation and AI agents for component-level diagnosis and remediation of failed or degraded hardware.
- Partner with data center operations to build tooling and AI agents for managing critical environments.
- Develop post-repair validation and testing tools, including burn-in, PyTorch, and NVIDIA NCCL, to ensure system stability and performance.
- Own deployment, monitoring, and operational support for developed tooling to maximize GPU fleet availability and performance.
- Develop automation and operational tooling for facilities management, power, and direct liquid-cooling hardware systems.
- Identify problems, rapidly develop scalable solutions, and ship them.
- Set technical direction for specific projects and execute independently or collaboratively.

## Requirements
- Software engineering experience.
- Expertise in distributed systems, reliability, and cloud platforms.
- Experience with Kubernetes, infrastructure as code, and Google Cloud.
- Proficiency in at least one programming language: Go, Python, Java, or Rust.
- Strong analytical and problem-solving skills.
- Excellent communication and collaboration skills.
- Ability to work independently and assist with critical or complex technical initiatives.

## Nice-to-Haves
- Experience with Temporal and Kubernetes.
- Experience working directly with hardware vendors.
- Experience operating large-scale GPU fleets or hyperscale data center environments.

## Compensation and Benefits
- Compensation range of **$215,000–$260,000 plus bonus**.
- Restricted Stock Units included in all offers.
- Health insurance options including HDHP and PPO, vision, and dental coverage.
- Employer HSA contributions.
- Paid parental leave.
- Paid life insurance and short- and long-term disability coverage.
- Teladoc.
- 401(k) with a 100% match up to 4% of salary.
- Paid time off and holidays.
- Cell phone and tuition reimbursement.
- Calm app subscription.
- MetLife Legal.
- Company-paid commuter benefit of $300 per month.

## Similar jobs

- [Staff Site Reliability Engineer - Site Experience](https://hotfix.jobs/jobs/d8fa80d2-f6f6-4a20-8485-60d2648f8daf) - Reddit - San Francisco, CA - $217k – $304k/yr
- [Staff Site Reliability Engineer, Ads](https://hotfix.jobs/jobs/91c6ec8f-96a4-469b-ad89-c0a0aeb77e26) - Reddit - Remote - $217k – $304k/yr
- [Staff Software Engineer, Observability](https://hotfix.jobs/jobs/002663cc-6743-4d3d-9bd8-a1e5c012ed86) - Reddit - Remote - $217k – $304k/yr
- [Staff Software Engineer, Service Tools](https://hotfix.jobs/jobs/7329e2b5-45c2-4ff7-8438-ca34edaee978) - Airbnb - Remote - $212k – $265k/yr
- [Staff Software Engineer, Traffic](https://hotfix.jobs/jobs/c3dbadf4-fe1e-4cb3-b5f0-4e1aad1aea73) - Temporal - Remote - $212k – $280k/yr

**Apply:** https://hotfix.jobs/jobs/ee5e02a2-cdbb-433e-88e7-c91608e76e93
**Canonical:** https://hotfix.jobs/jobs/ee5e02a2-cdbb-433e-88e7-c91608e76e93