Skip to content
Cerebras SystemsCerebras SystemsSan Francisco, CA

Site Reliability Engineer - Ops & Automation

Operates and automates production infrastructure for a high-scale AI inference service. The role requires production Kubernetes experience, Python or Go proficiency, observability expertise, and a focus on reliability, automation, and reducing operational toil.

Salary not listed
On-site5+ YOEDevOps / SRE

About the role

Responsibilities

  • Execute releases, capacity changes, and cluster upgrades while robust continuous delivery pipelines and self-service capabilities are developed.
  • Contribute to self-service continuous delivery pipelines using Kubernetes, Bazel, Prometheus, Grafana, InfluxDB, Python, and Go.
  • Build reusable automation and internal developer tools to minimize operational toil and cross-team friction.
  • Develop and extend telemetry, observability, and alerting solutions for reliability at scale.
  • Collaborate with Cluster Operations and development teams to identify high-impact automation opportunities.
  • Contribute to reliability practices, including SLOs, post-mortems, and capacity planning.

Requirements

  • 2–4+ years of experience in site reliability engineering with a strong operations or automation focus.
  • Production Kubernetes experience.
  • Strong Python or Go skills for building tools and automation.
  • Proficiency with Prometheus, Grafana, and observability-driven workflows.
  • Ability to measure and communicate impact through reliability metrics, operational toil, and velocity gains.

Nice to Have

  • Hands-on GitOps expertise with Argo CD, Flux, or equivalent tools.
  • Experience building continuous delivery pipelines.
  • Experience with Bazel or similar build systems.
  • Familiarity with capacity planning and on-premises or multi-datacenter environments.

Compensation and Benefits

  • The role does not require 24/7 on-call rotations.
  • Immediate ownership of production systems, direct mentorship from experienced engineers, and collaboration with Staff SREs.
  • Cerebras offers a non-corporate work culture, startup vitality, opportunities to work on a breakthrough AI platform and AI supercomputer, and support for continuous learning, growth, and inclusion.

Skills

KubernetesBazelPrometheusGrafanainfluxdbPythonGoGitOpsargo cdfluxcontinuous deliveryObservabilityslosCapacity Planning

Similar roles

DevOps / SRE jobs
Quindar

Site Reliability Engineer, US Gov

QuindarArvada, CO

Build and operate secure, highly available cloud infrastructure for mission-critical government and space systems across AWS GovCloud and C2E environments. The role requires Kubernetes, Terraform, Python, observability, networking, compliance, and an active U.S. security clearance.

160k – 200k/yrHybrid3+ YOEDevOps / SRE
OpenAI

Network Engineer

OpenAISan Francisco, CA

Designs, operates, and improves secure enterprise networks spanning offices, campuses, cloud environments, and connectivity services. The role combines architecture, production operations, troubleshooting, observability, security, and infrastructure automation.

293k – 385k/yrHybridDevOps / SRE
Applied Intuition

Build & Release Engineer

Applied IntuitionSunnyvale, CA

Owns software release workflows, dependency updates, artifact management, CI/CD pipelines, and an internal release portal. The role requires at least three years of software development experience, strong coding skills, and hands-on expertise with Git, CI/CD, and artifact repositories.

118k – 200k/yrOn-site3+ YOEDevOps / SRE
Anthropic

Software Engineer, Infrastructure, Interpretability

AnthropicSan Francisco, CA +1

Build secure, scalable infrastructure, data systems, compute tooling, and developer experiences for Anthropic’s Interpretability research team. The role partners closely with researchers, security, and platform teams and requires strong programming and infrastructure experience.

320k – 485k/yrHybridDevOps / SRE
Baseten

Software Engineer - Continuous Delivery

BasetenSan Francisco, CA +1

Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.

165k – 330k/yrHybridDevOps / SRE