# Site Reliability Engineer - Ops & Automation

**Company:** [Cerebras Systems](https://hotfix.jobs/companies/cerebras-systems)
**Location:** San Francisco, CA, Toronto, Canada
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** Kubernetes, Bazel, Prometheus, Grafana, influxdb, Python, Go, GitOps, argo cd, flux, continuous delivery, Observability, slos, Capacity Planning
**Posted:** 2025-10-14

> Operates and automates production infrastructure for a high-scale AI inference service. The role requires production Kubernetes experience, Python or Go proficiency, observability expertise, and a focus on reliability, automation, and reducing operational toil.

## Job Description

## Responsibilities
- Execute releases, capacity changes, and cluster upgrades while robust continuous delivery pipelines and self-service capabilities are developed.
- Contribute to self-service continuous delivery pipelines using Kubernetes, Bazel, Prometheus, Grafana, InfluxDB, Python, and Go.
- Build reusable automation and internal developer tools to minimize operational toil and cross-team friction.
- Develop and extend telemetry, observability, and alerting solutions for reliability at scale.
- Collaborate with Cluster Operations and development teams to identify high-impact automation opportunities.
- Contribute to reliability practices, including SLOs, post-mortems, and capacity planning.

## Requirements
- 2–4+ years of experience in site reliability engineering with a strong operations or automation focus.
- Production Kubernetes experience.
- Strong Python or Go skills for building tools and automation.
- Proficiency with Prometheus, Grafana, and observability-driven workflows.
- Ability to measure and communicate impact through reliability metrics, operational toil, and velocity gains.

## Nice to Have
- Hands-on GitOps expertise with Argo CD, Flux, or equivalent tools.
- Experience building continuous delivery pipelines.
- Experience with Bazel or similar build systems.
- Familiarity with capacity planning and on-premises or multi-datacenter environments.

## Compensation and Benefits
- The role does not require 24/7 on-call rotations.
- Immediate ownership of production systems, direct mentorship from experienced engineers, and collaboration with Staff SREs.
- Cerebras offers a non-corporate work culture, startup vitality, opportunities to work on a breakthrough AI platform and AI supercomputer, and support for continuous learning, growth, and inclusion.

## Similar roles

- [Site Reliability Engineer, US Gov](https://hotfix.jobs/jobs/ed7221ee-a969-4d2a-97cf-765f75ec65e3) - Quindar - Arvada, CO - $160k – $200k/yr
- [Network Engineer](https://hotfix.jobs/jobs/d6a557ff-247e-49d3-980e-b2ec395a4083) - OpenAI - San Francisco, CA - $293k – $385k/yr
- [Build & Release Engineer](https://hotfix.jobs/jobs/c41c87c5-6007-46e9-8648-635789c39b4a) - Applied Intuition - Sunnyvale, CA - $118k – $200k/yr
- [Software Engineer, Infrastructure, Interpretability](https://hotfix.jobs/jobs/b167b8e0-6f36-44a7-92b2-7803a2869b2a) - Anthropic - San Francisco, CA - $320k – $485k/yr
- [Software Engineer - Continuous Delivery](https://hotfix.jobs/jobs/0ceaa410-df08-45d6-80b3-68328c769297) - Baseten - San Francisco, CA - $165k – $330k/yr

**Apply:** https://hotfix.jobs/jobs/b89334fe-52d7-412e-9ca2-c499552fd275
**Canonical:** https://hotfix.jobs/jobs/b89334fe-52d7-412e-9ca2-c499552fd275