Build and operate distributed software for Cerebras wafer-scale AI clusters, including provisioning, orchestration, scheduling, monitoring, failure handling, and upgrade workflows. The role requires strong distributed-systems development experience and proficiency in Go, Python, Bash, Kubernetes, Prometheus, and Grafana.
Salary not listed
RemoteDevOps / SRE
About the role
Responsibilities
Automate bare-metal configuration of networking, operating systems, and application software in large clusters of Cerebras WSEs, servers, and switches.
Build push-button workflows for cluster upgrades, downgrades, and security patching, with key metrics to minimize downtime.
Develop an orchestration and scheduler system for resource allocation, job submission, and placement in a multi-user cluster environment.
Support both on-premises and cloud deployment and operations.
Build robust systems for monitoring, detecting, and handling failures across cluster resources, including high availability.
Develop cluster and job monitoring, visualization, and alerting capabilities.
Build user-facing tools to monitor job status and collect metrics.
Build administrator-facing tools to manage and operate large clusters.
Requirements
Strong track record in software architecture, system design, and development.
Build and operate secure, highly available cloud infrastructure for mission-critical government and space systems across AWS GovCloud and C2E environments. The role requires Kubernetes, Terraform, Python, observability, networking, compliance, and an active U.S. security clearance.
160k – 200k/yrHybrid3+ YOEDevOps / SRE
Network Engineer
OpenAISan Francisco, CA
Designs, operates, and improves secure enterprise networks spanning offices, campuses, cloud environments, and connectivity services. The role combines architecture, production operations, troubleshooting, observability, security, and infrastructure automation.
293k – 385k/yrHybridDevOps / SRE
Build & Release Engineer
Applied IntuitionSunnyvale, CA
Owns software release workflows, dependency updates, artifact management, CI/CD pipelines, and an internal release portal. The role requires at least three years of software development experience, strong coding skills, and hands-on expertise with Git, CI/CD, and artifact repositories.
Build secure, scalable infrastructure, data systems, compute tooling, and developer experiences for Anthropic’s Interpretability research team. The role partners closely with researchers, security, and platform teams and requires strong programming and infrastructure experience.
320k – 485k/yrHybridDevOps / SRE
Software Engineer - Continuous Delivery
BasetenSan Francisco, CA +1
Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.