Skip to content
CrusoeCrusoeSan Francisco, CA

Senior Staff Software Engineer, DC Infrastructure

Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.

250k – 300k/yr
On-site7+ YOEDevOps / SRE

About the role

Responsibilities

  • Develop deep-level diagnostics and troubleshooting for hardware faults within GPU racks and high-density compute systems.
  • Build troubleshooting and automation tooling for NVIDIA A100, H200, GB200, B200, and AMD 350X/355X GPU platforms.
  • Develop automation and AI agents for component-level diagnosis and remediation of failed or degraded hardware.
  • Partner with data center operations to create tooling and AI agents for managing critical environments.
  • Develop post-repair validation and testing tools, including burn-in, PyTorch, and NVIDIA NCCL, to ensure system stability and performance.
  • Own deployment, monitoring, and operational support for developed tooling to maximize GPU fleet availability and performance.
  • Develop automation and operational tooling for facility power and direct liquid-cooling hardware systems.
  • Set technical direction for projects and execute scalable solutions.
  • Support teammates working on critical or complex technical initiatives.

Requirements

  • Software engineering experience.
  • Experience with distributed systems, reliability, and cloud platforms.
  • Experience with Kubernetes, infrastructure as code, and Google Cloud.
  • Proficiency in at least one of Go, Python, Java, or Rust.
  • Strong analytical, problem-solving, communication, and collaboration skills.
  • Ability to work independently and within a team.

Nice to Have

  • Experience with Temporal and Kubernetes.
  • Experience working directly with hardware vendors.
  • Experience operating large-scale GPU fleets or hyperscale data center environments.

Compensation and Benefits

  • Compensation range of $250,000–$300,000 plus bonus.
  • Restricted Stock Units included in all offers.
  • Health insurance options including HDHP and PPO, vision, and dental coverage for employees and dependents.
  • Employer HSA contributions.
  • Paid parental leave.
  • Paid life insurance and short- and long-term disability coverage.
  • Teladoc.
  • 401(k) with a 100% match up to 4% of salary.
  • Generous paid time off and holiday schedule.
  • Cell phone reimbursement.
  • Tuition reimbursement.
  • Calm app subscription.
  • MetLife Legal.
  • Company-paid commuter benefit of $300 per month.

Skills

GoPythonJavaRustDistributed SystemsKubernetesInfrastructure As CodeGCPTemporalPyTorchnvidia ncclgpu fleet operationsAI Agentsdirect liquid cooling

Similar roles

DevOps / SRE jobs
Crusoe

Senior Staff Deployment Automation Engineer

CrusoeSan Francisco, CA +2

Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.

250k – 300k/yrOn-site12+ YOEDevOps / SRE
Perplexity

Member of Technical Staff

PerplexitySan Francisco, CA +2

Owns a multi-cloud GPU infrastructure platform that enables training and inference workloads through self-service orchestration. The role requires deep Kubernetes, distributed systems, GPU networking, and systems programming experience.

250k – 485k/yrRemoteDevOps / SRE
Perplexity

Member of Technical Staff

PerplexitySan Francisco, CA +1

Hands-on technical role building AI-powered tools, infrastructure, and processes to accelerate engineering velocity and product delivery at an AI search company.

250k – 405k/yrHybrid5+ YOEDevOps / SRE
Together AI

Staff Engineer, Distributed Storage and HPC & AI Infrastructure

Together AISan Francisco, CA

Design and operate multi-petabyte distributed storage systems for large-scale AI training and inference, integrating parallel filesystems and building Kubernetes-native storage platforms.

250k – 300k/yrOn-site8+ YOEDevOps / SRE
Zoox

Staff Site Reliability Engineer

ZooxFoster City, CA

Zoox is seeking a Staff Site Reliability Engineer to lead source control, owning the technical strategy and roadmap for their Git-based monorepo. This role involves migrating from GitHub Enterprise to GitHub Cloud, building developer tooling, and partnering with various teams to enhance source control as a strategic asset.

250k – 300k/yrHybrid7+ YOEDevOps / SRE