Skip to content
CrusoeCrusoeSan Francisco, CA

Senior Staff Deployment Automation Engineer

Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.

250k – 300k/yr
On-site12+ YOEDevOps / SRE

About the role

Responsibilities

  • Own deployment and integration testing automation for bare-metal, on-premise systems across the AI Cloud stack.
  • Build CI/CD platforms that enable rapid testing, iteration, and deployment of low-level systems and applications.
  • Design and execute large-scale validation tests across multi-node virtualized clusters to verify GPU workload scaling and stability.
  • Maintain and scale bare-metal Linux configurations using GitLab, Ansible, AWX, osquery, and related tooling.
  • Create control applications for canary deployments, blue/green testing, and automated rollback on production systems.
  • Develop automation frameworks in Python or Go to provision, configure, and stress-test multi-node virtualized environments.
  • Build automated test suites with fio, stress-ng, and iperf to validate performance and CPU/GPU host isolation.

Requirements

  • 12+ years of professional experience performing comparable responsibilities independently.
  • Bachelor's or master's degree in Computer Science, Electrical Engineering, or a related technical field.
  • Experience building and deploying automated integration testing for AI cloud environments, from low-level Linux systems through distributed control planes.
  • Working knowledge of Kubernetes, Docker, Terraform, and PostgreSQL.
  • Extensive knowledge of CI/CD pipelines and GitLab tooling across multiple datacenters.
  • Experience with one or more configuration management systems, such as Ansible, Puppet, Chef, or SaltStack.
  • Advanced Python and/or Bash proficiency for complex cluster-wide automation.
  • Knowledge of Linux kernel internals, including PCIe topology, VFIO, HugePages, and IOMMU.
  • Familiarity with NVIDIA CUDA/NCCL and/or AMD ROCm/RCCL in multi-node environments.
  • Strong understanding of RDMA, RoCE, and InfiniBand in virtualized systems.

Nice-to-haves

  • Experience with MNNVL or specialized AI fabric architectures.
  • Familiarity with hardware debugging tools and performance profilers such as NVIDIA Nsight and AMD Omniperf.
  • Knowledge of GPU container orchestration, including Kubernetes device plugins.

Compensation and Benefits

  • Compensation of up to $250,000–$300,000, plus bonus.
  • Restricted Stock Units included in offers.
  • Paid time off, holidays, and leave programs.
  • Health, dental, and vision insurance.
  • Employer HSA contributions.
  • Paid parental leave, life insurance, and short- and long-term disability coverage.
  • Professional development and tuition reimbursement.
  • Mental health and wellness support.
  • Commuter benefits and cell phone stipend.
  • 401(k) plan with company match up to 4% of salary.
  • Volunteer time off, global travel insurance, emergency assistance, daily meals allowance, and location-specific programs.

Skills

PythonBashGoKubernetesDockerTerraformPostgresGitLabAnsibleLinuxpcievfioCUDAncclrdma

Similar roles

DevOps / SRE jobs
Crusoe

Senior Staff Software Engineer, DC Infrastructure

CrusoeSan Francisco, CA +1

Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.

250k – 300k/yrOn-site7+ YOEDevOps / SRE
Perplexity

Member of Technical Staff

PerplexitySan Francisco, CA +2

Owns a multi-cloud GPU infrastructure platform that enables training and inference workloads through self-service orchestration. The role requires deep Kubernetes, distributed systems, GPU networking, and systems programming experience.

250k – 485k/yrRemoteDevOps / SRE
Perplexity

Member of Technical Staff

PerplexitySan Francisco, CA +1

Hands-on technical role building AI-powered tools, infrastructure, and processes to accelerate engineering velocity and product delivery at an AI search company.

250k – 405k/yrHybrid5+ YOEDevOps / SRE
Together AI

Staff Engineer, Distributed Storage and HPC & AI Infrastructure

Together AISan Francisco, CA

Design and operate multi-petabyte distributed storage systems for large-scale AI training and inference, integrating parallel filesystems and building Kubernetes-native storage platforms.

250k – 300k/yrOn-site8+ YOEDevOps / SRE
Zoox

Staff Site Reliability Engineer

ZooxFoster City, CA

Zoox is seeking a Staff Site Reliability Engineer to lead source control, owning the technical strategy and roadmap for their Git-based monorepo. This role involves migrating from GitHub Enterprise to GitHub Cloud, building developer tooling, and partnering with various teams to enhance source control as a strategic asset.

250k – 300k/yrHybrid7+ YOEDevOps / SRE