# Software Engineer, Cluster Deployment

**Company:** [Cerebras Systems](https://hotfix.jobs/companies/cerebras-systems)
**Location:** Sunnyvale, CA
**Role:** DevOps / SRE
**Experience:** 0+ years
**Skills:** Python, Bash, Linux, Git, Terraform, Ansible, Kubernetes, Networking, Prometheus, Grafana
**Posted:** 2026-07-23

> Build and maintain automation tooling for large-scale AI compute cluster deployments, turning bare-metal infrastructure into repeatable, pushbutton workflows using Python, Ansible, Terraform, Kubernetes and observability tools. Ideal for new graduates or early-career engineers seeking hands-on production infrastructure experience.

## Job Description

## Responsibilities
- Develop and maintain automation for deployment workflows, including provisioning, configuration, validation, and operational handoff.
- Turn manual deployment steps into tested, repeatable pushbutton workflows.
- Participate in hands-on cluster deployments to build practical debugging and operational expertise.
- Troubleshoot issues across Linux systems, bare-metal servers, networking, storage, Kubernetes, and connectivity.
- Contribute to infrastructure-as-code and GitOps workflows using tools such as Terraform, Ansible, pull requests, and code review.
- Add health checks, observability, dashboards, and validation logic to improve deployment reliability.
- Partner with networking, infrastructure, security, and operations teams to deliver secure and reproducible data center deployments.

## Basic Qualifications
- Strong fundamentals in Python and Bash, with the ability to write scripts and small programs.
- Basic Linux experience, including command-line usage, processes, filesystems, and disk troubleshooting.
- Working knowledge of Git, including branching, commits, pull requests, and code review.
- CS, ECE, or related technical degree, or equivalent practical experience.
- Curiosity, strong problem-solving instincts, and willingness to work hands-on with real infrastructure.

## Preferred Qualifications
- Networking fundamentals, including VLANs and routing basics; exposure to BGP, switch configuration, or Arista EOS automation is a plus.
- Kubernetes experience or familiarity.
- Infrastructure-as-code and GitOps experience, including Terraform, Ansible, and PR-based change control.
- Bare-metal provisioning concepts such as PXE, DHCP, iPXE, Redfish, IPMI, and BMC management.
- Observability experience with Prometheus or Grafana.
- API design and client-server architecture.
- Automation side projects or open-source contributions.

## Similar roles

- [Site Reliability Engineer II](https://hotfix.jobs/jobs/e9632016-5af4-4276-b34d-b6dc80c872cf) - Instacart - Remote - $133k – $169k/yr
- [Software Engineer, Infrastructure (2-8 YOE)](https://hotfix.jobs/jobs/cc716b3a-a9f1-478f-90f4-4f36d971cf33) - Airtable - San Francisco, CA - $148k – $250k/yr
- [Associate Infrastructure Engineer](https://hotfix.jobs/jobs/08d10862-c584-4779-9996-09875ba9850a) - Webflow - Remote - $140k – $190k/yr
- [Operations Engineer](https://hotfix.jobs/jobs/26207eb4-a57c-4702-b762-2c0fd6399396) - xAI - Memphis, TN
- [Network Deployment Engineer](https://hotfix.jobs/jobs/aa90e15c-41ab-4c19-a306-16d96672e748) - Cloudflare - Atlanta, GA - $126k – $173k/yr

**Apply:** https://hotfix.jobs/jobs/df9b7cb2-588d-4dd1-b0c2-8bdcc284fd11
**Canonical:** https://hotfix.jobs/jobs/df9b7cb2-588d-4dd1-b0c2-8bdcc284fd11