# Principal Engineer, AI Inference Reliability

**Company:** [Cerebras Systems](https://hotfix.jobs/companies/cerebras-systems)
**Location:** Unspecified
**Role:** DevOps / SRE
**Experience:** 7+ years
**Skills:** Python, C++, Go, Rust, Distributed Systems, slo design, sli design, sla design, Incident Response, chaos testing, load simulation, fault injection, Observability, multi-region deployments, AI Infrastructure
**Posted:** 2025-10-29

> Leads reliability strategy and hands-on implementation for a large-scale, low-latency AI inference service. The role requires 7+ years in backend, infrastructure, or reliability engineering, strong backend programming skills, and deep expertise in distributed-system reliability.

## Job Description

## Responsibilities
- Define and drive reliability strategy, including establishing SLOs and aligning engineering teams.
- Design and implement reliability mechanisms for fault detection, graceful degradation, failover, throttling, and recovery across multiple regions and data centers.
- Lead large-scale incident management, including postmortems, root-cause analysis, and prevention loops.
- Architect for reliability and observability by influencing system design for redundancy, durability, and debuggability.
- Develop internal tooling and frameworks for chaos testing, load simulation, and distributed fault injection.
- Collaborate with software, infrastructure, and hardware teams to embed reliability across the inference service.
- Build dashboards and alerts to monitor service health and communicate actionable reliability metrics.
- Mentor engineers and establish best practices for designing, testing, and operating reliable large-scale systems.

## Requirements
- Bachelor's or master's degree in computer science or a related field.
- 7+ years of experience in backend, infrastructure, or reliability engineering for large-scale distributed systems.
- Strong programming skills in at least one backend programming language, such as Python, C++, Go, or Rust.
- Deep experience with reliability principles, including SLO/SLI/SLA design, incident response, and postmortem culture.
- Excellent communication and cross-functional leadership skills.

## Nice-to-haves
- Experience building large-scale AI infrastructure systems.

## Benefits
- Work on a breakthrough AI platform beyond the constraints of GPU-based systems.
- Publish and open-source cutting-edge AI research.
- Work on one of the world's fastest AI supercomputers.
- Enjoy job stability with startup vitality.
- Participate in a non-corporate work culture that respects individual beliefs.

## Similar roles

- [Principal Operations Engineer, Mechanical](https://hotfix.jobs/jobs/1033a3a6-548c-4699-8e58-e9cb70cf32cd) - Fluidstack - Remote - $150k – $250k/yr
- [Principal Systems Engineer, DevTools](https://hotfix.jobs/jobs/138f8113-11f6-41cb-bcff-1d5236e63542) - Cloudflare - Atlanta, GA - $200k – $281k/yr
- [Principal Operations Engineer, Controls](https://hotfix.jobs/jobs/9c821ce8-5edc-4d88-8572-e74fac66a8b0) - Fluidstack - Remote - $150k – $250k/yr
- [Principal Operations Engineer, Electrical](https://hotfix.jobs/jobs/56295001-a8c6-4fb8-ba2b-1c6cbf808ad8) - Fluidstack - Remote - $150k – $250k/yr
- [Principal Classified Systems Architect, Okta Federal](https://hotfix.jobs/jobs/84afc3e2-4b38-48c9-bd76-066efb2f6fc5) - Okta - Washington, DC - $224k – $308k/yr

**Apply:** https://hotfix.jobs/jobs/2d675d6e-753b-40db-a7dd-da27b8903033
**Canonical:** https://hotfix.jobs/jobs/2d675d6e-753b-40db-a7dd-da27b8903033