Skip to content
Cerebras SystemsCerebras SystemsUnited States

Principal Engineer, AI Inference Reliability

Leads reliability strategy and hands-on implementation for a large-scale, low-latency AI inference service. The role requires 7+ years in backend, infrastructure, or reliability engineering, strong backend programming skills, and deep expertise in distributed-system reliability.

Salary not listed
On-site7+ YOEDevOps / SRE

About the role

Responsibilities

  • Define and drive reliability strategy, including establishing SLOs and aligning engineering teams.
  • Design and implement reliability mechanisms for fault detection, graceful degradation, failover, throttling, and recovery across multiple regions and data centers.
  • Lead large-scale incident management, including postmortems, root-cause analysis, and prevention loops.
  • Architect for reliability and observability by influencing system design for redundancy, durability, and debuggability.
  • Develop internal tooling and frameworks for chaos testing, load simulation, and distributed fault injection.
  • Collaborate with software, infrastructure, and hardware teams to embed reliability across the inference service.
  • Build dashboards and alerts to monitor service health and communicate actionable reliability metrics.
  • Mentor engineers and establish best practices for designing, testing, and operating reliable large-scale systems.

Requirements

  • Bachelor's or master's degree in computer science or a related field.
  • 7+ years of experience in backend, infrastructure, or reliability engineering for large-scale distributed systems.
  • Strong programming skills in at least one backend programming language, such as Python, C++, Go, or Rust.
  • Deep experience with reliability principles, including SLO/SLI/SLA design, incident response, and postmortem culture.
  • Excellent communication and cross-functional leadership skills.

Nice-to-haves

  • Experience building large-scale AI infrastructure systems.

Benefits

  • Work on a breakthrough AI platform beyond the constraints of GPU-based systems.
  • Publish and open-source cutting-edge AI research.
  • Work on one of the world's fastest AI supercomputers.
  • Enjoy job stability with startup vitality.
  • Participate in a non-corporate work culture that respects individual beliefs.

Skills

PythonC++GoRustDistributed Systemsslo designsli designsla designIncident Responsechaos testingload simulationfault injectionObservabilitymulti-region deploymentsAI Infrastructure

Similar roles

DevOps / SRE jobs
Fluidstack

Principal Operations Engineer, Mechanical

FluidstackUnited States

As a Principal Operations Engineer, Mechanical, you will be the senior technical authority for mechanical and cooling infrastructure across hyperscale AI data centers. You will lead site assessments, drive operational readiness, review designs, and ensure precision execution of critical systems.

150k – 250k/yrRemote10+ YOEDevOps / SRE
Cloudflare

Principal Systems Engineer, DevTools

CloudflareAtlanta, GA +3

Build and operate AI-powered developer tools, internal MCP integrations, and platform capabilities across the engineering organization. The role requires strong coding and debugging skills, Kubernetes operations experience, and the ability to lead projects, improve developer experience, and mentor teammates.

200k – 281k/yrHybrid7+ YOEDevOps / SRE
Fluidstack

Principal Operations Engineer, Controls

FluidstackUnited States

As Principal Operations Engineer, Controls, you will be the senior technical authority for operational building automation and control systems across hyperscale AI data centers. You will lead site assessments, drive technical readiness, review designs, and ensure the continuous improvement of control systems.

150k – 250k/yrRemote10+ YOEDevOps / SRE
Fluidstack

Principal Operations Engineer, Electrical

FluidstackUnited States

Fluidstack is seeking a Principal Operations Engineer, Electrical to be the senior technical authority for electrical infrastructure across their hyperscale AI data center portfolio. This role involves leading site assessments, driving technical readiness, reviewing designs, and feeding operational learnings back into the design and manufacturing organization.

150k – 250k/yrRemote10+ YOEDevOps / SRE
Okta

Principal Classified Systems Architect, Okta Federal

OktaWashington, DC

Leads architecture of air-gapped classified (SIPR/JWICS) developer platforms for DoD compliance, designs IaC for disconnected ops, integrates hardened tools like Big Bang/Iron Bank, and ensures secure, scalable Kubernetes infrastructure. Requires 12+ years experience with 5+ in classified DoD environments.

224k – 308k/yrOn-site12+ YOEDevOps / SRE