Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.
250k – 300k/yr
On-site12+ YOEDevOps / SRE
About the role
Responsibilities
Own deployment and integration testing automation for bare-metal, on-premise systems across the AI Cloud stack.
Build CI/CD platforms that enable rapid testing, iteration, and deployment of low-level systems and applications.
Design and execute large-scale validation tests across multi-node virtualized clusters to verify GPU workload scaling and stability.
Maintain and scale bare-metal Linux configurations using GitLab, Ansible, AWX, osquery, and related tooling.
Create control applications for canary deployments, blue/green testing, and automated rollback on production systems.
Develop automation frameworks in Python or Go to provision, configure, and stress-test multi-node virtualized environments.
Build automated test suites with fio, stress-ng, and iperf to validate performance and CPU/GPU host isolation.
Requirements
12+ years of professional experience performing comparable responsibilities independently.
Bachelor's or master's degree in Computer Science, Electrical Engineering, or a related technical field.
Experience building and deploying automated integration testing for AI cloud environments, from low-level Linux systems through distributed control planes.
Working knowledge of Kubernetes, Docker, Terraform, and PostgreSQL.
Extensive knowledge of CI/CD pipelines and GitLab tooling across multiple datacenters.
Experience with one or more configuration management systems, such as Ansible, Puppet, Chef, or SaltStack.
Advanced Python and/or Bash proficiency for complex cluster-wide automation.
Knowledge of Linux kernel internals, including PCIe topology, VFIO, HugePages, and IOMMU.
Familiarity with NVIDIA CUDA/NCCL and/or AMD ROCm/RCCL in multi-node environments.
Strong understanding of RDMA, RoCE, and InfiniBand in virtualized systems.
Nice-to-haves
Experience with MNNVL or specialized AI fabric architectures.
Familiarity with hardware debugging tools and performance profilers such as NVIDIA Nsight and AMD Omniperf.
Knowledge of GPU container orchestration, including Kubernetes device plugins.
Compensation and Benefits
Compensation of up to $250,000–$300,000, plus bonus.
Restricted Stock Units included in offers.
Paid time off, holidays, and leave programs.
Health, dental, and vision insurance.
Employer HSA contributions.
Paid parental leave, life insurance, and short- and long-term disability coverage.
Professional development and tuition reimbursement.
Mental health and wellness support.
Commuter benefits and cell phone stipend.
401(k) plan with company match up to 4% of salary.
Volunteer time off, global travel insurance, emergency assistance, daily meals allowance, and location-specific programs.
Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.
250k – 300k/yrOn-site7+ YOEDevOps / SRE
Member of Technical Staff
PerplexitySan Francisco, CA +2
Owns a multi-cloud GPU infrastructure platform that enables training and inference workloads through self-service orchestration. The role requires deep Kubernetes, distributed systems, GPU networking, and systems programming experience.
250k – 485k/yrRemoteDevOps / SRE
Member of Technical Staff
PerplexitySan Francisco, CA +1
Hands-on technical role building AI-powered tools, infrastructure, and processes to accelerate engineering velocity and product delivery at an AI search company.
250k – 405k/yrHybrid5+ YOEDevOps / SRE
Staff Engineer, Distributed Storage and HPC & AI Infrastructure
Together AISan Francisco, CA
Design and operate multi-petabyte distributed storage systems for large-scale AI training and inference, integrating parallel filesystems and building Kubernetes-native storage platforms.
250k – 300k/yrOn-site8+ YOEDevOps / SRE
Staff Site Reliability Engineer
ZooxFoster City, CA
Zoox is seeking a Staff Site Reliability Engineer to lead source control, owning the technical strategy and roadmap for their Git-based monorepo. This role involves migrating from GitHub Enterprise to GitHub Cloud, building developer tooling, and partnering with various teams to enhance source control as a strategic asset.