# AI Infrastructure Operations Engineer

**Company:** [Cerebras Systems](https://hotfix.jobs/companies/cerebras-systems)
**Location:** Unspecified
**Role:** DevOps / SRE
**Experience:** 0+ years
**Skills:** Linux, server hardware, Networking, hardware telemetry, monitoring systems, Incident Response, log collection, cluster hardware
**Posted:** 2026-02-09

> Entry-level SiteOps engineer supporting deployment, validation, monitoring, and first-line troubleshooting of AI clusters in data center environments. Requires a relevant engineering degree or equivalent experience, familiarity with server hardware, networking, and Linux, and readiness to work hands-on in data centers.

## Job Description

## Responsibilities
- Assist with deployment and bring-up of CS-X systems, cluster servers, and networking hardware.
  - Execute power-on sequencing, readiness checks, and validation tests.
- Monitor hardware telemetry, alerts, and dashboards.
- Perform first-line troubleshooting and structured escalation.
- Collect logs, telemetry, and observations during incidents.

## Incident Support & Tooling
- Participate in incident response under senior engineer guidance.
- Use existing monitoring, telemetry, and incident tracking tools.
- Provide feedback on tooling and process gaps.

## Learning & Development
- Build working knowledge of Cerebras system architecture.
- Learn cluster hardware and networking fundamentals.
- Shadow senior engineers during complex debugging.
- Progress toward independent ownership of defined workflows.

## Requirements
- Bachelor's degree in a relevant engineering field or equivalent experience.
- 0–3 years of experience in hardware operations, systems engineering, or data center environments.
- Basic familiarity with server hardware, networking fundamentals, and Linux systems.

## Nice-to-Haves
- Internship or early-career experience in data center or hardware lab environments.
- Exposure to monitoring or telemetry systems.
- Comfort working in data centers.

## Success Measures
- Consistent and correct execution of hardware bring-up procedures.
- Early identification and escalation of issues.
- Improved documentation quality.
- Clear progression toward more independent operational responsibility.

## Similar roles

- [Site Reliability Engineer II](https://hotfix.jobs/jobs/e9632016-5af4-4276-b34d-b6dc80c872cf) - Instacart - Remote - $133k – $169k/yr
- [Software Engineer, Infrastructure (2-8 YOE)](https://hotfix.jobs/jobs/cc716b3a-a9f1-478f-90f4-4f36d971cf33) - Airtable - San Francisco, CA - $148k – $250k/yr
- [Associate Infrastructure Engineer](https://hotfix.jobs/jobs/08d10862-c584-4779-9996-09875ba9850a) - Webflow - Remote - $140k – $190k/yr
- [Operations Engineer](https://hotfix.jobs/jobs/26207eb4-a57c-4702-b762-2c0fd6399396) - xAI - Memphis, TN
- [Network Deployment Engineer](https://hotfix.jobs/jobs/aa90e15c-41ab-4c19-a306-16d96672e748) - Cloudflare - Atlanta, GA - $126k – $173k/yr

**Apply:** https://hotfix.jobs/jobs/2d5feb00-978b-4a89-98f6-77a7d3ab939e
**Canonical:** https://hotfix.jobs/jobs/2d5feb00-978b-4a89-98f6-77a7d3ab939e