# Engineering Manager, Kernel Reliability

**Company:** [Cerebras Systems](https://hotfix.jobs/companies/cerebras-systems)
**Location:** Unspecified
**Role:** Engineering Management
**Experience:** 6+ years
**Skills:** Parallel Programming, Distributed Systems, Message Passing, Multicore Computing, Gpu Computing, Embedded Systems, Debuggers, Core Dumps, Code Sanitizers, Failure Analysis, Computer Architecture, Multithreading, Networking, Monitoring, Incident Response
**Posted:** 2026-01-08

> Leads a hands-on Kernel Reliability team responsible for improving the reliability, diagnostics, and failure analysis of advanced compute clusters and production services. The role requires software engineering expertise, distributed-systems debugging experience, and proven engineering leadership.

## Job Description

## Responsibilities
- Provide hands-on technical leadership, owning the technical vision and roadmap for kernel-centric reliability of internal and customer-facing systems.
- Assist System and Cluster Operations teams in reducing system and service downtime after failures through tooling and manual intervention for failure analysis and diagnostics.
- Work with the Debug Team to enhance debugging tools and speed failure analysis.
- Collaborate with software teams to improve the software stack, including kernels, for on-field debugging and failure analysis.
- Work with ASIC and hardware architecture teams to co-design next-generation architectures with reliability and ease of debugging in mind.
- Lead, mentor, and grow a high-caliber engineering team while fostering technical excellence and rapid execution.

## Requirements
- 6+ years of software engineering experience.
- 3+ years leading teams in software or hardware reliability, debugging, diagnostics, failure analysis, or related fields.
- Expertise in parallel and distributed programming, including message passing, multicore, GPU, and embedded systems.
- Experience developing or using debugging and diagnostic tools, including debuggers, core dump handling, and code sanitizers.
- Experience debugging distributed and parallel applications, including deadlocks, livelocks, and race conditions.
- Deep understanding of computer architecture, including instruction pipelining, multithreading, and networking.
- Strong background in monitoring and reliability engineering, including incident response and post-mortem analysis.
- Ability to recruit and retain high-performing teams, mentor engineers, and partner cross-functionally to deliver customer-facing products.

## Compensation & Benefits
- Opportunity to build a breakthrough AI platform beyond the constraints of GPUs.
- Ability to publish and open-source cutting-edge AI research.
- Opportunity to work on one of the fastest AI supercomputers in the world.
- Job stability with startup vitality.
- Non-corporate work culture that respects individual beliefs.
- Continuous learning, growth, and support.

## Similar jobs

- [Senior Engineering Manager — Network Connectivity](https://hotfix.jobs/jobs/a18d25e4-dee6-4f67-bda7-49d1b744ff86) - Cloudflare - Austin, TX - €89k – €122k/yr
- [Senior Engineering Manager, SDET](https://hotfix.jobs/jobs/63a066f6-ecda-4c63-ad27-7d1b60754de1) - Cribl - Remote - $225k – $270k/yr
- [Engineering Manager of Managers, Service Infrastructure](https://hotfix.jobs/jobs/74ad6d7e-1ee0-418e-a688-5aba5cf65a53) - Stripe - Seattle, WA
- [Engineering Manager, App Traffic](https://hotfix.jobs/jobs/b52c3c30-68bf-4cef-b721-80047ae6da93) - Databricks - Mountain View, CA - $190k – $238k/yr
- [Senior Engineering Manager, Mortgage](https://hotfix.jobs/jobs/504bc007-9071-49ae-9cc4-f579dd230a96) - Checkr - San Francisco, CA - $269k – $316k/yr

**Apply:** https://hotfix.jobs/jobs/a7ed06ad-a093-4468-b4ea-0aaa15e4da24
**Canonical:** https://hotfix.jobs/jobs/a7ed06ad-a093-4468-b4ea-0aaa15e4da24