Skip to content
Cerebras SystemsCerebras SystemsUnited States

Distributed Software Engineer

Build and operate distributed software for Cerebras wafer-scale AI clusters, including provisioning, orchestration, scheduling, monitoring, failure handling, and upgrade workflows. The role requires strong distributed-systems development experience and proficiency in Go, Python, Bash, Kubernetes, Prometheus, and Grafana.

Salary not listed
RemoteDevOps / SRE

About the role

Responsibilities

  • Automate bare-metal configuration of networking, operating systems, and application software in large clusters of Cerebras WSEs, servers, and switches.
  • Build push-button workflows for cluster upgrades, downgrades, and security patching, with key metrics to minimize downtime.
  • Develop an orchestration and scheduler system for resource allocation, job submission, and placement in a multi-user cluster environment.
  • Support both on-premises and cloud deployment and operations.
  • Build robust systems for monitoring, detecting, and handling failures across cluster resources, including high availability.
  • Develop cluster and job monitoring, visualization, and alerting capabilities.
  • Build user-facing tools to monitor job status and collect metrics.
  • Build administrator-facing tools to manage and operate large clusters.

Requirements

  • Strong track record in software architecture, system design, and development.
  • Strong experience developing distributed cluster systems.
  • Strong understanding of the Kubernetes software ecosystem, Prometheus, and Grafana.
  • Strong development skills in Go, Python, and Bash.
  • Strong debugging skills with distributed systems.
  • Strong ability to develop tests for new features and prevent regressions.

Skills

KubernetesPrometheusGrafanaGoPythonBashDistributed Systemscluster orchestrationcluster schedulingbare-metal provisioningMonitoringhigh availabilityautomated testing

Similar roles

DevOps / SRE jobs
Quindar

Site Reliability Engineer, US Gov

QuindarArvada, CO

Build and operate secure, highly available cloud infrastructure for mission-critical government and space systems across AWS GovCloud and C2E environments. The role requires Kubernetes, Terraform, Python, observability, networking, compliance, and an active U.S. security clearance.

160k – 200k/yrHybrid3+ YOEDevOps / SRE
OpenAI

Network Engineer

OpenAISan Francisco, CA

Designs, operates, and improves secure enterprise networks spanning offices, campuses, cloud environments, and connectivity services. The role combines architecture, production operations, troubleshooting, observability, security, and infrastructure automation.

293k – 385k/yrHybridDevOps / SRE
Applied Intuition

Build & Release Engineer

Applied IntuitionSunnyvale, CA

Owns software release workflows, dependency updates, artifact management, CI/CD pipelines, and an internal release portal. The role requires at least three years of software development experience, strong coding skills, and hands-on expertise with Git, CI/CD, and artifact repositories.

118k – 200k/yrOn-site3+ YOEDevOps / SRE
Anthropic

Software Engineer, Infrastructure, Interpretability

AnthropicSan Francisco, CA +1

Build secure, scalable infrastructure, data systems, compute tooling, and developer experiences for Anthropic’s Interpretability research team. The role partners closely with researchers, security, and platform teams and requires strong programming and infrastructure experience.

320k – 485k/yrHybridDevOps / SRE
Baseten

Software Engineer - Continuous Delivery

BasetenSan Francisco, CA +1

Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.

165k – 330k/yrHybridDevOps / SRE