# Staff Software Engineer

**Company:** [Crusoe](https://hotfix.jobs/companies/crusoe)
**Location:** San Francisco, CA
**Role:** Embedded Engineering
**Salary:** $215k – $260k/yr
**Experience:** 7+ years
**Skills:** Go, gpu architectures, nvidia a100, nvidia h200, nvidia gb200, nvidia b200, amd 350x, amd 355x, InfiniBand, nvlink, roce, Linux, ubuntu, rocky linux, centos
**Posted:** 2026-07-23

> Build and maintain software for advanced diagnosis, troubleshooting, and repair of Crusoe's large-scale GPU compute clusters (NVIDIA and AMD). Requires strong Golang coding, deep GPU/hardware expertise, Linux proficiency, and experience in production data center environments.

## Job Description

## What You’ll Be Doing
- Automate deep-level diagnosis and troubleshooting of hardware faults within GPU racks and high-density compute systems.
- Develop software to troubleshoot and support GPU platforms including NVIDIA A100, H200, GB200, B200, B300 and AMD 350X / 355X.
- Execute component-level diagnosis and remediation for failed or degraded hardware.
- Partner with data center operations to manage and perform field-replaceable unit (FRU) repairs for GPUs, power supplies, cooling systems, interconnects, and networking hardware.
- Conduct post-repair validation, burn-in testing, torch testing, and NVIDIA NCCL testing to ensure system stability and performance.
- Implement and execute preventative maintenance procedures to improve fleet reliability and extend hardware lifespan.
- Perform firmware and BIOS upgrades across the GPU fleet.
- Maintain detailed documentation of maintenance activities, failures, and resolutions in ticketing and asset management systems.
- Develop and update standard operating procedures (SOPs) for troubleshooting, repair, and validation workflows.
- Collaborate with engineering, software, and data center operations teams to identify root causes of systemic failures and implement preventative solutions.
- Participate in a rotating infrastructure on-call schedule (about one week every 4–6 weeks) with daytime coverage and handoff to the Europe team.

## What You’ll Bring to the Team
- Ability to code in Golang
- Proven experience diagnosing and repairing high-density, rack-mounted compute hardware in production environments.
- Deep understanding of GPU architectures and hands-on experience with GPU-based systems.
- Experience supporting NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X series platforms.
- Familiarity with high-speed interconnects such as InfiniBand, NVLink, and RDMA over Converged Ethernet (RoCE).
- Strong Linux experience (Ubuntu, Rocky Linux, CentOS) using the command line for diagnostics and testing.
- Proficiency with GPU and system diagnostic tools such as NVIDIA DCGM and NVIDIA field diagnostic utilities.
- Experience working with enterprise server hardware, power delivery, and cooling systems.
- Strong analytical and problem-solving skills.
- Excellent communication and collaboration skills.
- Ability to work independently in a fast-paced data center or operations environment.

## Nice to Have
- Technical certification or Associate’s/Bachelor’s degree in Electrical Engineering, Computer Science, or a related field or demonstrated experience.
- Experience working directly with hardware vendors and escalations.
- Background in large-scale GPU fleet operations or hyperscale data center environments.

## Benefits
- Hybrid work schedule
- Industry competitive pay
- Restricted Stock Units in a fast growing, well-funded technology company
- Health insurance package options that include HDHP and PPO, vision, and dental for you and your dependents
- Employer contributions to HSA accounts
- Paid Parental Leave
- Paid life insurance, short-term and long-term disability
- Teladoc
- 401(k) with a 100% match up to 4% of salary
- Generous paid time off and holiday schedule
- Cell phone reimbursement
- Tuition reimbursement
- Subscription to the Calm app
- MetLife Legal
- Company paid commuter benefit; $300 per pay period

## Compensation
Compensation will be paid in the range of $215,000 - $260,000. Restricted Stock Units are included in all offers. Compensation to be determined by the applicants knowledge, education, and abilities, as well as internal equity and alignment with market data.

## Similar roles

- [Senior / Staff Connectivity Integration Lead](https://hotfix.jobs/jobs/927394e1-aac0-4ada-b39c-e104c373eea8) - Zoox - Foster City, CA - $220k – $275k/yr
- [Senior Staff Engineer, Flight Controls - X-BAT (R4812)](https://hotfix.jobs/jobs/4848979b-3869-4d74-9000-8910bb76bdf6) - Shield AI - Dallas, TX - $220k – $340k/yr
- [Senior Staff Engineer, Software - Autonomous Aircraft Integration (R4983)](https://hotfix.jobs/jobs/21adca5a-476e-4b6c-a8cc-297f3b246ddf) - Shield AI - Washington, DC - $221k – $331k/yr
- [Senior Staff Engineer, Autonomy - Tactical Behaviors (R4073)](https://hotfix.jobs/jobs/3c403ec6-8a00-4bb9-8816-71b26ec17943) - Shield AI - Washington, DC - $221k – $331k/yr
- [Staff Embedded Software Engineer](https://hotfix.jobs/jobs/bac0a112-d82c-4ca0-b162-8e41837367a4) - Cyngn - Mountain View, CA - $205k – $220k/yr

**Apply:** https://hotfix.jobs/jobs/efda99bf-0f63-4484-851e-913681c5fbbc
**Canonical:** https://hotfix.jobs/jobs/efda99bf-0f63-4484-851e-913681c5fbbc