Engineering Manager, Kernel Reliability
Leads a hands-on Kernel Reliability team responsible for improving the reliability, diagnostics, and failure analysis of advanced compute clusters and production services. The role requires software engineering expertise, distributed-systems debugging experience, and proven engineering leadership.
About the job
Responsibilities
- Provide hands-on technical leadership, owning the technical vision and roadmap for kernel-centric reliability of internal and customer-facing systems.
- Assist System and Cluster Operations teams in reducing system and service downtime after failures through tooling and manual intervention for failure analysis and diagnostics.
- Work with the Debug Team to enhance debugging tools and speed failure analysis.
- Collaborate with software teams to improve the software stack, including kernels, for on-field debugging and failure analysis.
- Work with ASIC and hardware architecture teams to co-design next-generation architectures with reliability and ease of debugging in mind.
- Lead, mentor, and grow a high-caliber engineering team while fostering technical excellence and rapid execution.
Requirements
- 6+ years of software engineering experience.
- 3+ years leading teams in software or hardware reliability, debugging, diagnostics, failure analysis, or related fields.
- Expertise in parallel and distributed programming, including message passing, multicore, GPU, and embedded systems.
- Experience developing or using debugging and diagnostic tools, including debuggers, core dump handling, and code sanitizers.
- Experience debugging distributed and parallel applications, including deadlocks, livelocks, and race conditions.
- Deep understanding of computer architecture, including instruction pipelining, multithreading, and networking.
- Strong background in monitoring and reliability engineering, including incident response and post-mortem analysis.
- Ability to recruit and retain high-performing teams, mentor engineers, and partner cross-functionally to deliver customer-facing products.
Compensation & Benefits
- Opportunity to build a breakthrough AI platform beyond the constraints of GPUs.
- Ability to publish and open-source cutting-edge AI research.
- Opportunity to work on one of the fastest AI supercomputers in the world.
- Job stability with startup vitality.
- Non-corporate work culture that respects individual beliefs.
- Continuous learning, growth, and support.
Skills
Parallel Programming, Distributed Systems, Message Passing, Multicore Computing, Gpu Computing, Embedded Systems, Debuggers, Core Dumps, Code Sanitizers, Failure Analysis, Computer Architecture, Multithreading, Networking, Monitoring, Incident Response
Similar jobs
Engineering Management jobsLeads a distributed team building and operating network egress infrastructure, owning technical vision, delivery, reliability, and production operations. Requires at least five years of engineering management experience plus expertise or interest in networking and distributed systems.
Leads a remote-first team of SDETs responsible for product quality, test strategy, automation, and release reliability across SaaS and customer-managed environments. Requires 10+ years of industry experience, technical leadership, and strong expertise in testing, CI/CD, observability, and quality metrics.
Leads multiple service-infrastructure engineering teams and managers, shaping architecture, developer platforms, reliability, and cross-functional delivery. Requires substantial management experience, including managing managers, critical distributed systems, incident response, and geographically distributed teams.
Leads and develops the App Traffic engineering team building reliable, scalable service-mesh and networking infrastructure across multiple clouds. Requires 9+ years of software engineering experience, including engineering leadership and distributed-systems or infrastructure expertise.
Leads the mortgage engineering organization, owning platform architecture, delivery, business-line outcomes, and team development. Requires senior engineering management experience, extensive software engineering experience, large-team leadership, business ownership, and expertise in scalable systems and AI.