# Senior Staff Software Engineer, DC Infrastructure

**Company:** [Crusoe](https://hotfix.jobs/companies/crusoe)
**Location:** San Francisco, CA, Sunnyvale, CA
**Role:** DevOps / SRE
**Salary:** $250k – $300k/yr
**Experience:** 7+ years
**Skills:** Go, Python, Java, Rust, Distributed Systems, Kubernetes, Infrastructure As Code, GCP, Temporal, PyTorch, nvidia nccl, gpu fleet operations, AI Agents, direct liquid cooling
**Posted:** 2026-08-12

> Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.

## Job Description

## Responsibilities
- Develop deep-level diagnostics and troubleshooting for hardware faults within GPU racks and high-density compute systems.
- Build troubleshooting and automation tooling for NVIDIA A100, H200, GB200, B200, and AMD 350X/355X GPU platforms.
- Develop automation and AI agents for component-level diagnosis and remediation of failed or degraded hardware.
- Partner with data center operations to create tooling and AI agents for managing critical environments.
- Develop post-repair validation and testing tools, including burn-in, PyTorch, and NVIDIA NCCL, to ensure system stability and performance.
- Own deployment, monitoring, and operational support for developed tooling to maximize GPU fleet availability and performance.
- Develop automation and operational tooling for facility power and direct liquid-cooling hardware systems.
- Set technical direction for projects and execute scalable solutions.
- Support teammates working on critical or complex technical initiatives.

## Requirements
- Software engineering experience.
- Experience with distributed systems, reliability, and cloud platforms.
- Experience with Kubernetes, infrastructure as code, and Google Cloud.
- Proficiency in at least one of Go, Python, Java, or Rust.
- Strong analytical, problem-solving, communication, and collaboration skills.
- Ability to work independently and within a team.

## Nice to Have
- Experience with Temporal and Kubernetes.
- Experience working directly with hardware vendors.
- Experience operating large-scale GPU fleets or hyperscale data center environments.

## Compensation and Benefits
- Compensation range of **$250,000–$300,000 plus bonus**.
- Restricted Stock Units included in all offers.
- Health insurance options including HDHP and PPO, vision, and dental coverage for employees and dependents.
- Employer HSA contributions.
- Paid parental leave.
- Paid life insurance and short- and long-term disability coverage.
- Teladoc.
- 401(k) with a 100% match up to 4% of salary.
- Generous paid time off and holiday schedule.
- Cell phone reimbursement.
- Tuition reimbursement.
- Calm app subscription.
- MetLife Legal.
- Company-paid commuter benefit of $300 per month.

## Similar roles

- [Senior Staff Deployment Automation Engineer](https://hotfix.jobs/jobs/e7c5a05a-7667-45ba-8f89-00349f1e9aaa) - Crusoe - San Francisco, CA - $250k – $300k/yr
- [Member of Technical Staff](https://hotfix.jobs/jobs/e2ff6e88-b443-42b7-87ae-33997af4c66f) - Perplexity - Remote - $250k – $485k/yr
- [Member of Technical Staff](https://hotfix.jobs/jobs/ca5edb5f-58ba-4e0c-a075-25c3fcc61007) - Perplexity - San Francisco, CA - $250k – $405k/yr
- [Staff Engineer, Distributed Storage and HPC & AI Infrastructure](https://hotfix.jobs/jobs/d0fb38e7-169a-4b2d-abb6-6f845d3f381f) - Together AI - San Francisco, CA - $250k – $300k/yr
- [Staff Site Reliability Engineer](https://hotfix.jobs/jobs/e2420bd8-af00-4aa8-9bcb-34da7d97403d) - Zoox - Foster City, CA - $250k – $300k/yr

**Apply:** https://hotfix.jobs/jobs/5e51b5cc-872b-4556-a065-f5f95a363fad
**Canonical:** https://hotfix.jobs/jobs/5e51b5cc-872b-4556-a065-f5f95a363fad