# Senior Staff Deployment Automation Engineer

**Company:** [Crusoe](https://hotfix.jobs/companies/crusoe)
**Location:** San Francisco, CA, Sunnyvale, CA, Seattle, WA
**Role:** DevOps / SRE
**Salary:** $250k – $300k/yr
**Experience:** 12+ years
**Skills:** Python, Bash, Go, Kubernetes, Docker, Terraform, Postgres, GitLab, Ansible, Linux, pcie, vfio, CUDA, nccl, rdma
**Posted:** 2026-08-13

> Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.

## Job Description

## Responsibilities
- Own deployment and integration testing automation for bare-metal, on-premise systems across the AI Cloud stack.
- Build CI/CD platforms that enable rapid testing, iteration, and deployment of low-level systems and applications.
- Design and execute large-scale validation tests across multi-node virtualized clusters to verify GPU workload scaling and stability.
- Maintain and scale bare-metal Linux configurations using GitLab, Ansible, AWX, osquery, and related tooling.
- Create control applications for canary deployments, blue/green testing, and automated rollback on production systems.
- Develop automation frameworks in Python or Go to provision, configure, and stress-test multi-node virtualized environments.
- Build automated test suites with fio, stress-ng, and iperf to validate performance and CPU/GPU host isolation.

## Requirements
- **12+ years of professional experience** performing comparable responsibilities independently.
- Bachelor's or master's degree in Computer Science, Electrical Engineering, or a related technical field.
- Experience building and deploying automated integration testing for AI cloud environments, from low-level Linux systems through distributed control planes.
- Working knowledge of Kubernetes, Docker, Terraform, and PostgreSQL.
- Extensive knowledge of CI/CD pipelines and GitLab tooling across multiple datacenters.
- Experience with one or more configuration management systems, such as Ansible, Puppet, Chef, or SaltStack.
- Advanced Python and/or Bash proficiency for complex cluster-wide automation.
- Knowledge of Linux kernel internals, including PCIe topology, VFIO, HugePages, and IOMMU.
- Familiarity with NVIDIA CUDA/NCCL and/or AMD ROCm/RCCL in multi-node environments.
- Strong understanding of RDMA, RoCE, and InfiniBand in virtualized systems.

## Nice-to-haves
- Experience with MNNVL or specialized AI fabric architectures.
- Familiarity with hardware debugging tools and performance profilers such as NVIDIA Nsight and AMD Omniperf.
- Knowledge of GPU container orchestration, including Kubernetes device plugins.

## Compensation and Benefits
- Compensation of up to **$250,000–$300,000**, plus bonus.
- Restricted Stock Units included in offers.
- Paid time off, holidays, and leave programs.
- Health, dental, and vision insurance.
- Employer HSA contributions.
- Paid parental leave, life insurance, and short- and long-term disability coverage.
- Professional development and tuition reimbursement.
- Mental health and wellness support.
- Commuter benefits and cell phone stipend.
- 401(k) plan with company match up to 4% of salary.
- Volunteer time off, global travel insurance, emergency assistance, daily meals allowance, and location-specific programs.

## Similar roles

- [Senior Staff Software Engineer, DC Infrastructure](https://hotfix.jobs/jobs/5e51b5cc-872b-4556-a065-f5f95a363fad) - Crusoe - San Francisco, CA - $250k – $300k/yr
- [Member of Technical Staff](https://hotfix.jobs/jobs/e2ff6e88-b443-42b7-87ae-33997af4c66f) - Perplexity - Remote - $250k – $485k/yr
- [Member of Technical Staff](https://hotfix.jobs/jobs/ca5edb5f-58ba-4e0c-a075-25c3fcc61007) - Perplexity - San Francisco, CA - $250k – $405k/yr
- [Staff Engineer, Distributed Storage and HPC & AI Infrastructure](https://hotfix.jobs/jobs/d0fb38e7-169a-4b2d-abb6-6f845d3f381f) - Together AI - San Francisco, CA - $250k – $300k/yr
- [Staff Site Reliability Engineer](https://hotfix.jobs/jobs/e2420bd8-af00-4aa8-9bcb-34da7d97403d) - Zoox - Foster City, CA - $250k – $300k/yr

**Apply:** https://hotfix.jobs/jobs/e7c5a05a-7667-45ba-8f89-00349f1e9aaa
**Canonical:** https://hotfix.jobs/jobs/e7c5a05a-7667-45ba-8f89-00349f1e9aaa