# Distributed Software Engineer

**Company:** [Cerebras Systems](https://hotfix.jobs/companies/cerebras-systems)
**Location:** Sunnyvale, CA
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** Go, Python, Kubernetes, Custom Resource Definitions, Kubernetes Operators, gRPC, Linux, Networking, Prometheus, Grafana, Promql, RBAC, Ebpf, Ceph, Rdma
**Posted:** 2026-09-08

> Build and operate distributed infrastructure software that automates, schedules, observes, and repairs large-scale AI compute clusters. The role requires 5+ years of infrastructure or distributed-systems experience, strong Go and Python skills, and deep Kubernetes expertise.

## Job Description

## Responsibilities
- Build declarative, CRD-driven automation for bare-metal networking, operating systems, and application software across clusters of Cerebras systems, servers, and switches.
- Develop push-button cluster installation, upgrades, and security patching with canary gates and defined downtime budgets.
- Build Kubernetes operators for large-scale inference workload scheduling, including resource locks, priority queues, network topology, and health-aware placement.
- Develop gRPC control-plane services, authorization, admission webhooks, and quota policies for a multi-tenant fleet.
- Build metrics and log pipelines with exporters for wafer-scale systems, servers, and network fabrics using Prometheus and Grafana.
- Implement failure detection, highly available control planes, automated recovery, and CLIs, APIs, and MCP gateways for fleet management.

## Requirements
- 5+ years of experience building and operating production distributed systems or infrastructure software.
- Production-quality Go and Python programming skills.
- Deep Kubernetes experience, including controllers, operators, CRDs, reconciliation semantics, informer caches, admission webhooks, and RBAC.
- Strong debugging skills across distributed systems, Linux, and networking.
- Practical Prometheus and Grafana experience, including PromQL, exporter design, cardinality management, alerting, and SLOs.
- Ability to work independently in a fast-moving, incompletely documented environment and drive cross-team initiatives.
- Demonstrated use of AI tools in software engineering, with the ability to evaluate and verify generated output.

## Nice to Have
- Bare-metal or HPC fleet operations.
- Scheduler internals.
- RDMA/RoCE and eBPF networking.
- Ceph or NVMe-oF.
- etcd and high-availability upgrades.
- Inference-serving stacks.

## Similar jobs

- [Software Engineer, Infrastructure](https://hotfix.jobs/jobs/585da47a-e02d-4c63-9816-248a2faa9b5b) - Granica - Remote
- [Capacity Ops Engineer](https://hotfix.jobs/jobs/f1904714-7dd3-4ee3-9e7a-e4fcf52083bd) - Baseten - San Francisco, CA - $170k – $230k/yr
- [IT Security and Automation Engineer](https://hotfix.jobs/jobs/604b87b5-13a2-4bba-88b2-f7d0fbbad141) - Teleport - Remote - $149k – $258k/yr
- [Electrical Field Engineer - Data Center](https://hotfix.jobs/jobs/6bfa0e4c-9ccf-438a-b65e-cd4c6297762c) - Crusoe - Remote - $196k – $235k/yr
- [Software Engineer, Cloud Infrastructure](https://hotfix.jobs/jobs/949677d6-6d57-49e8-acf8-017a14790019) - Beacon AI - San Carlos, CA - $135k – $260k/yr

**Apply:** https://hotfix.jobs/jobs/58d3b54d-348d-4797-9849-3905f4846800
**Canonical:** https://hotfix.jobs/jobs/58d3b54d-348d-4797-9849-3905f4846800