Skip to content

Latest DevOps / SRE jobs at Cerebras Systems

Search
Location
21 jobs

Job results

Cerebras Systems

Cluster Operations Software Engineer

Cerebras SystemsSunnyvale, CA

Operates and develops software for large-scale AI compute clusters, improving reliability, capacity, monitoring, and incident response. Requires 6–8 years of complex infrastructure experience, strong Python and Go skills, distributed-systems expertise, and participation in 24/7 on-call support.

Salary not listedHybrid6+ YOEDevOps / SRE
Cerebras Systems

AI Inference Core - Senior SW Engineer for Platform & DevOps

Cerebras SystemsUnited States

Build and operate the platform layer behind Cerebras engineering infrastructure, including CI/CD, Kubernetes, deployment automation, cloud and on-premises systems, developer environments, and observability. The role requires 5+ years of infrastructure or software engineering experience and strong debugging and systems fundamentals.

Salary not listedHybrid5+ YOEDevOps / SRE
Cerebras Systems

Software Engineer, Cluster Deployment

Cerebras SystemsSunnyvale, CA

Build and maintain automation tooling for large-scale AI compute cluster deployments, turning bare-metal infrastructure into repeatable, pushbutton workflows using Python, Ansible, Terraform, Kubernetes and observability tools. Ideal for new graduates or early-career engineers seeking hands-on production infrastructure experience.

Salary not listedOn-siteEntry levelDevOps / SRE
Cerebras Systems

Infrastructure Engineer

Cerebras SystemsSunnyvale, CA

Infrastructure Engineer responsible for hands-on installation, provisioning, maintenance, and troubleshooting of high-performance on-premise server hardware, Linux systems, and high-speed networking (100G/400G) in a data center environment. Requires 3+ years experience with Linux admin, x86 hardware, and network configuration.

Salary not listedOn-site3+ YOEDevOps / SRE
Cerebras Systems

Software Engineer

Cerebras SystemsSunnyvale, CA

Build and maintain CI/CD pipelines, artifact management, cloud infrastructure, and developer productivity tooling at Cerebras to accelerate AI hardware and software engineering. Requires 2-5 years DevOps/infrastructure experience, Kubernetes, AWS, and strong troubleshooting skills.

Salary not listedOn-site2+ YOEDevOps / SRE
Cerebras Systems

Cloud Infrastructure Engineer

Cerebras SystemsSunnyvale, CA

Design, build, and operate secure cloud infrastructure and identity platforms in AWS and on-prem data centers. Implement IAM, automation with Terraform/Python/Go, security controls for AI systems, and Zero Trust principles while participating in on-call.

Salary not listedOn-site5+ YOEDevOps / SRE
Cerebras Systems

Principal Site Reliability Engineer

Cerebras SystemsSunnyvale, CA

Principal SRE to architect self-service reliability platforms, capacity orchestration, and production control planes for Cerebras' ultra-high-speed AI inference infrastructure at massive scale. Requires 15+ years in SRE/platform engineering with deep large-scale fleet and observability experience.

Salary not listedOn-site15+ YOEDevOps / SRE
Cerebras Systems

Software Engineer, Inference Platform

Cerebras SystemsSunnyvale, CA

Software engineer building and operating the orchestration layer for a globally distributed, high-performance AI inference platform on custom wafer-scale hardware.

Salary not listedOn-site3+ YOEDevOps / SRE
Cerebras Systems

Staff Software Engineer, Inference Platform

Cerebras SystemsSunnyvale, CA

Hands-on technical lead building and operating the orchestration layer for a globally distributed, high-performance AI inference platform on custom wafer-scale hardware.

Salary not listedOn-site8+ YOEDevOps / SRE
Cerebras Systems

Network Engineer

Cerebras SystemsSunnyvale, CA

Design and operate large-scale AI/HPC cluster network fabrics. Architect front-end datacenter interconnects, build automation and observability tooling, and debug complex distributed networking issues.

Salary not listedOn-site5+ YOEDevOps / SRE
Cerebras Systems

Member of Technical Staff (Software Engineer)

Cerebras SystemsSunnyvale, CA

Develops and optimizes Kubernetes-based infrastructure for high-performance AI inference services, including deployment, scaling, debugging, and integration with ML workflows. Requires Master's in CS and 1+ year experience with Docker, Kubernetes, Python, and related tools.

170k – 175k/yrRemoteDevOps / SRE
Cerebras Systems

Sr. Member of Technical Staff

Cerebras SystemsSunnyvale, CA

Develops resilient, high-availability software for AI inference on AWS, including deployment workflows, container orchestration with Docker/Kubernetes, monitoring, and debugging. Requires Master's in CS and 18 months experience with AWS services, IaC tools, and Python.

230k – 250k/yrHybridDevOps / SRE
Cerebras Systems

Senior WAN Network Engineer

Cerebras SystemsSunnyvale, CA

Designs, implements, and optimizes global WAN networks using leased lines, dark fiber, and advanced routing protocols like BGP for low-latency, high-availability connectivity. Requires 6+ years experience, bachelor's degree, CCIE/JNCIE certs, and expertise in automation tools like Python/Ansible/Terraform.

Salary not listedOn-site6+ YOEDevOps / SRE
Cerebras Systems

Staff Software Engineer, Inference Cloud

Cerebras SystemsSunnyvale, CA

Staff engineer owns architecture of Inference Cloud Platform, building distributed systems for high-QPS AI workloads with focus on availability, latency, reliability, and global scale. Requires 8+ years experience in large-scale cloud systems and backend languages like Go, C++, Python.

Salary not listedOn-site8+ YOEDevOps / SRE
Cerebras Systems

Principal Engineer, Inference Cloud

Cerebras SystemsSunnyvale, CA

Principal Engineer leads Inference Cloud Platform, defining architecture for multi-region, high-QPS AI inference systems. Focuses on reliability, performance optimization, production code, and cross-team technical strategy. Requires 10+ years in distributed systems.

Salary not listedOn-site10+ YOEDevOps / SRE
Cerebras Systems

Staff Site Reliability Engineer – Automation and Platform

Cerebras SystemsSunnyvale, CA

Leads automation and platform engineering for ultra-reliable AI inference infrastructure, architecting self-service GitOps pipelines, observability, and tooling to eliminate toil across datacenters. Requires 8+ years SRE experience with large-scale clusters and tools like Argo CD and Prometheus.

Salary not listedRemote8+ YOEDevOps / SRE
Cerebras Systems

Design Validation Test - Lead/Principal Engineer

Cerebras SystemsSunnyvale, CA

Leads end-to-end Design Validation Test (DVT) for complex electrical boards and systems, including power delivery, high-speed I/O validation, debug, and root-cause analysis. Requires 8+ years experience in hardware validation, strong EE skills, and lab equipment proficiency.

175k – 275k/yrOn-site8+ YOEDevOps / SRE
Cerebras Systems

AI Infrastructure Operations Engineer

Cerebras SystemsUnited States

Entry-level SiteOps engineer supporting deployment, validation, monitoring, and first-line troubleshooting of AI clusters in data center environments. Requires a relevant engineering degree or equivalent experience, familiarity with server hardware, networking, and Linux, and readiness to work hands-on in data centers.

Salary not listedOn-siteEntry levelDevOps / SRE
Cerebras Systems

Distributed Software Engineer

Cerebras SystemsUnited States

Build and operate distributed software for Cerebras wafer-scale AI clusters, including provisioning, orchestration, scheduling, monitoring, failure handling, and upgrade workflows. The role requires strong distributed-systems development experience and proficiency in Go, Python, Bash, Kubernetes, Prometheus, and Grafana.

Salary not listedRemoteDevOps / SRE
Cerebras Systems

Principal Engineer, AI Inference Reliability

Cerebras SystemsUnited States

Leads reliability strategy and hands-on implementation for a large-scale, low-latency AI inference service. The role requires 7+ years in backend, infrastructure, or reliability engineering, strong backend programming skills, and deep expertise in distributed-system reliability.

Salary not listedOn-site7+ YOEDevOps / SRE
Cerebras Systems

Site Reliability Engineer - Ops & Automation

Cerebras SystemsSan Francisco, CA +1

Operates and automates production infrastructure for a high-scale AI inference service. The role requires production Kubernetes experience, Python or Go proficiency, observability expertise, and a focus on reliability, automation, and reducing operational toil.

Salary not listedOn-site5+ YOEDevOps / SRE