Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.
250k – 300k/yr
On-site7+ YOEDevOps / SRE
About the role
Responsibilities
Develop deep-level diagnostics and troubleshooting for hardware faults within GPU racks and high-density compute systems.
Build troubleshooting and automation tooling for NVIDIA A100, H200, GB200, B200, and AMD 350X/355X GPU platforms.
Develop automation and AI agents for component-level diagnosis and remediation of failed or degraded hardware.
Partner with data center operations to create tooling and AI agents for managing critical environments.
Develop post-repair validation and testing tools, including burn-in, PyTorch, and NVIDIA NCCL, to ensure system stability and performance.
Own deployment, monitoring, and operational support for developed tooling to maximize GPU fleet availability and performance.
Develop automation and operational tooling for facility power and direct liquid-cooling hardware systems.
Set technical direction for projects and execute scalable solutions.
Support teammates working on critical or complex technical initiatives.
Requirements
Software engineering experience.
Experience with distributed systems, reliability, and cloud platforms.
Experience with Kubernetes, infrastructure as code, and Google Cloud.
Proficiency in at least one of Go, Python, Java, or Rust.
Strong analytical, problem-solving, communication, and collaboration skills.
Ability to work independently and within a team.
Nice to Have
Experience with Temporal and Kubernetes.
Experience working directly with hardware vendors.
Experience operating large-scale GPU fleets or hyperscale data center environments.
Compensation and Benefits
Compensation range of $250,000–$300,000 plus bonus.
Restricted Stock Units included in all offers.
Health insurance options including HDHP and PPO, vision, and dental coverage for employees and dependents.
Employer HSA contributions.
Paid parental leave.
Paid life insurance and short- and long-term disability coverage.
Teladoc.
401(k) with a 100% match up to 4% of salary.
Generous paid time off and holiday schedule.
Cell phone reimbursement.
Tuition reimbursement.
Calm app subscription.
MetLife Legal.
Company-paid commuter benefit of $300 per month.
Skills
GoPythonJavaRustDistributed SystemsKubernetesInfrastructure As CodeGCPTemporalPyTorchnvidia ncclgpu fleet operationsAI Agentsdirect liquid cooling
Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.
250k – 300k/yrOn-site12+ YOEDevOps / SRE
Member of Technical Staff
PerplexitySan Francisco, CA +2
Owns a multi-cloud GPU infrastructure platform that enables training and inference workloads through self-service orchestration. The role requires deep Kubernetes, distributed systems, GPU networking, and systems programming experience.
250k – 485k/yrRemoteDevOps / SRE
Member of Technical Staff
PerplexitySan Francisco, CA +1
Hands-on technical role building AI-powered tools, infrastructure, and processes to accelerate engineering velocity and product delivery at an AI search company.
250k – 405k/yrHybrid5+ YOEDevOps / SRE
Staff Engineer, Distributed Storage and HPC & AI Infrastructure
Together AISan Francisco, CA
Design and operate multi-petabyte distributed storage systems for large-scale AI training and inference, integrating parallel filesystems and building Kubernetes-native storage platforms.
250k – 300k/yrOn-site8+ YOEDevOps / SRE
Staff Site Reliability Engineer
ZooxFoster City, CA
Zoox is seeking a Staff Site Reliability Engineer to lead source control, owning the technical strategy and roadmap for their Git-based monorepo. This role involves migrating from GitHub Enterprise to GitHub Cloud, building developer tooling, and partnering with various teams to enhance source control as a strategic asset.