Skip to content

Engineering Manager, Kernel Reliability

Leads a hands-on Kernel Reliability team responsible for improving the reliability, diagnostics, and failure analysis of advanced compute clusters and production services. The role requires software engineering expertise, distributed-systems debugging experience, and proven engineering leadership.

About the job

Responsibilities

  • Provide hands-on technical leadership, owning the technical vision and roadmap for kernel-centric reliability of internal and customer-facing systems.
  • Assist System and Cluster Operations teams in reducing system and service downtime after failures through tooling and manual intervention for failure analysis and diagnostics.
  • Work with the Debug Team to enhance debugging tools and speed failure analysis.
  • Collaborate with software teams to improve the software stack, including kernels, for on-field debugging and failure analysis.
  • Work with ASIC and hardware architecture teams to co-design next-generation architectures with reliability and ease of debugging in mind.
  • Lead, mentor, and grow a high-caliber engineering team while fostering technical excellence and rapid execution.

Requirements

  • 6+ years of software engineering experience.
  • 3+ years leading teams in software or hardware reliability, debugging, diagnostics, failure analysis, or related fields.
  • Expertise in parallel and distributed programming, including message passing, multicore, GPU, and embedded systems.
  • Experience developing or using debugging and diagnostic tools, including debuggers, core dump handling, and code sanitizers.
  • Experience debugging distributed and parallel applications, including deadlocks, livelocks, and race conditions.
  • Deep understanding of computer architecture, including instruction pipelining, multithreading, and networking.
  • Strong background in monitoring and reliability engineering, including incident response and post-mortem analysis.
  • Ability to recruit and retain high-performing teams, mentor engineers, and partner cross-functionally to deliver customer-facing products.

Compensation & Benefits

  • Opportunity to build a breakthrough AI platform beyond the constraints of GPUs.
  • Ability to publish and open-source cutting-edge AI research.
  • Opportunity to work on one of the fastest AI supercomputers in the world.
  • Job stability with startup vitality.
  • Non-corporate work culture that respects individual beliefs.
  • Continuous learning, growth, and support.

Skills

Parallel Programming, Distributed Systems, Message Passing, Multicore Computing, Gpu Computing, Embedded Systems, Debuggers, Core Dumps, Code Sanitizers, Failure Analysis, Computer Architecture, Multithreading, Networking, Monitoring, Incident Response

Cloudflare

Cloudflare

Austin, TX
Senior Engineering Manager — Network Connectivity
€89k+/yrHybrid5+ YOEEngineering Management

Leads a distributed team building and operating network egress infrastructure, owning technical vision, delivery, reliability, and production operations. Requires at least five years of engineering management experience plus expertise or interest in networking and distributed systems.

Cribl

Cribl

United States

Senior Engineering Manager, SDET
$225k+/yrRemote10+ YOEEngineering Management

Leads a remote-first team of SDETs responsible for product quality, test strategy, automation, and release reliability across SaaS and customer-managed environments. Requires 10+ years of industry experience, technical leadership, and strong expertise in testing, CI/CD, observability, and quality metrics.

Stripe

Stripe

Seattle, WA
Engineering Manager of Managers, Service Infrastructure
No salary listedOn-site8+ YOEEngineering Management

Leads multiple service-infrastructure engineering teams and managers, shaping architecture, developer platforms, reliability, and cross-functional delivery. Requires substantial management experience, including managing managers, critical distributed systems, incident response, and geographically distributed teams.

Databricks

Databricks

Mountain View, CA

Engineering Manager, App Traffic
$190k+/yrOn-site9+ YOEEngineering Management

Leads and develops the App Traffic engineering team building reliable, scalable service-mesh and networking infrastructure across multiple clouds. Requires 9+ years of software engineering experience, including engineering leadership and distributed-systems or infrastructure expertise.

Checkr

Checkr

San Francisco, CA

Senior Engineering Manager, Mortgage
$269k+/yrHybrid8+ YOEEngineering Management

Leads the mortgage engineering organization, owning platform architecture, delivery, business-line outcomes, and team development. Requires senior engineering management experience, extensive software engineering experience, large-team leadership, business ownership, and expertise in scalable systems and AI.