Skip to content
Cerebras SystemsCerebras SystemsUnited States

AI Infrastructure Operations Engineer

Entry-level SiteOps engineer supporting deployment, validation, monitoring, and first-line troubleshooting of AI clusters in data center environments. Requires a relevant engineering degree or equivalent experience, familiarity with server hardware, networking, and Linux, and readiness to work hands-on in data centers.

Salary not listed
On-siteEntry levelDevOps / SRE

About the role

Responsibilities

  • Assist with deployment and bring-up of CS-X systems, cluster servers, and networking hardware.
    • Execute power-on sequencing, readiness checks, and validation tests.
  • Monitor hardware telemetry, alerts, and dashboards.
  • Perform first-line troubleshooting and structured escalation.
  • Collect logs, telemetry, and observations during incidents.

Incident Support & Tooling

  • Participate in incident response under senior engineer guidance.
  • Use existing monitoring, telemetry, and incident tracking tools.
  • Provide feedback on tooling and process gaps.

Learning & Development

  • Build working knowledge of Cerebras system architecture.
  • Learn cluster hardware and networking fundamentals.
  • Shadow senior engineers during complex debugging.
  • Progress toward independent ownership of defined workflows.

Requirements

  • Bachelor's degree in a relevant engineering field or equivalent experience.
  • 0–3 years of experience in hardware operations, systems engineering, or data center environments.
  • Basic familiarity with server hardware, networking fundamentals, and Linux systems.

Nice-to-Haves

  • Internship or early-career experience in data center or hardware lab environments.
  • Exposure to monitoring or telemetry systems.
  • Comfort working in data centers.

Success Measures

  • Consistent and correct execution of hardware bring-up procedures.
  • Early identification and escalation of issues.
  • Improved documentation quality.
  • Clear progression toward more independent operational responsibility.

Skills

Linuxserver hardwareNetworkinghardware telemetrymonitoring systemsIncident Responselog collectioncluster hardware

Similar roles

DevOps / SRE jobs
Instacart

Site Reliability Engineer II

InstacartUnited States

Supports the reliability and performance of large-scale systems through monitoring, incident response, automation, troubleshooting, and dependable deployments. The role requires 2–4 years of software engineering experience and familiarity with scripting, system administration, and cloud platforms is preferred.

133k – 169k/yrRemote2+ YOEDevOps / SRE
Airtable

Software Engineer, Infrastructure (2-8 YOE)

AirtableSan Francisco, CA +3

Backend engineers build and scale Airtable's infrastructure across teams like Base, Compute, Data, Storage, and Traffic. Requires 2-8 years experience in distributed systems, databases; CS degree; hybrid work in SF, NYC, Seattle, or LA areas.

148k – 250k/yrHybrid2+ YOEDevOps / SRE
Webflow

Associate Infrastructure Engineer

WebflowUnited States

Operates and improves Webflow’s production infrastructure, focusing on Kubernetes reliability, observability, cloud systems, infrastructure as code, and incident response. The role requires 2+ years debugging distributed systems and comfort with application code and on-call operations.

140k – 190k/yrRemote2+ YOEDevOps / SRE
xAI

Operations Engineer

xAIMemphis, TN +1

Improves facility operations through process standardization, maintenance and construction workflow optimization, operational dashboards, data analysis, and automation. Requires an engineering bachelor's degree, at least one year of operations or process-improvement experience, and Excel, SQL, or Python skills.

Salary not listedOn-site1+ YOEDevOps / SRE
Cloudflare

Network Deployment Engineer

CloudflareAtlanta, GA +3

Deploys and expands global datacenter and physical network infrastructure, coordinating contractors, vendors, installations, and operational processes. Requires at least two years of datacenter or Linux systems administration experience plus networking, configuration management, scripting, and project coordination skills.

126k – 173k/yrHybrid2+ YOEDevOps / SRE