Skip to content
BasetenBaseten

Engineering Manager, Runtime Fabric

Lead the Runtime Fabrics team building container runtime and storage layers purpose-built for AI inference workloads. Manage systems engineers, set technical direction for containerd/runc extensions, and drive open-source contributions.

About the job

Responsibilities

Team Leadership & Culture

  • Recruit, hire, and develop a high-performing team of systems engineers with deep container and Linux expertise
  • Foster a culture of technical rigor, open-source contribution, and continuous improvement
  • Provide regular coaching, feedback, and career development support to direct reports
  • Partner with engineering leadership to define the long-term vision and roadmap for container runtime and storage infrastructure

Technical Direction

  • Guide the team in extending and hardening containerd, runc, and related OCI ecosystem projects to meet the GPU-specific requirements of production AI inference, including startup performance, GPU device access, and multi-tenant isolation
  • Oversee the architecture and evolution of the Baseten Delivery Network: the tiered caching and weight delivery system that makes cold starts 2–3x faster and eliminates thundering herd failures during burst scaling events
  • Drive the expansion of BDN's architecture, currently focused on model weights, to container images, training checkpoints, and deployment artifacts
  • Provide technical oversight on GPU-aware isolation mechanisms for multi-tenant inference, including secure container runtimes, Linux namespace hardening, and longer-term micro-VM integration
  • Ensure the team maintains end-to-end ownership of the container startup performance path, from snapshotter initialization through weight delivery to first inference request
  • Champion the team's contributions back to the open-source containerd ecosystem alongside a team of core maintainers

Cross-Functional Partnership

  • Act as the primary advocate for Runtime Fabrics across the organization, ensuring upstream and downstream teams have the integration support they need
  • Collaborate with product and engineering stakeholders to prioritize investments based on business impact and infrastructure reliability
  • Communicate team progress, technical trade-offs, and architectural decisions clearly to leadership

Requirements

  • Proven experience managing and growing engineering teams in a systems, infrastructure, or low-level runtime context
  • Deep familiarity with the Linux container ecosystem: containerd, runc, OCI Runtime Spec, Linux namespaces, and cgroups, with the ability to engage credibly in code reviews and architectural discussions
  • Contributions to containerd/containerd, opencontainers/runc, google/gvisor, kata-containers/kata-containers, or closely related open-source projects
  • Strong systems programming background in Go and/or C/C++
  • Experience with distributed storage systems, content-addressable storage, or large-scale caching infrastructure
  • Understanding of how container images are structured, stored, and delivered at scale
  • Strong written and verbal communication skills, with the ability to influence without authority across teams

Nice to Have

  • Experience with GPU device access in containers: NVIDIA Container Toolkit, CDI (Container Device Interface), or GPU-aware scheduling
  • Familiarity with lazy-loading snapshotters (stargz, soci, EROFS/Nydus) or peer-to-peer image distribution
  • Experience with secure container runtimes (gVisor, Sysbox) or micro-VM technologies (Firecracker, Cloud Hypervisor)
  • Understanding of containerd's shim API (v2) and experience building custom shim implementations
  • Background in multi-tenant infrastructure or security-sensitive serving environments

Benefits

  • Competitive compensation, including meaningful equity
  • 100% coverage of medical, dental, and vision insurance for employee and dependents
  • Flexible PTO policy including company wide Winter Break
  • Paid parental leave
  • Fertility and family-building stipend through Carrot
  • Company-facilitated 401(k)

Skills

Containerd, Runc, Oci Runtime Spec, Linux Namespaces, Cgroups, Go, C, C++, Distributed Storage Systems, Content-Addressable Storage, Large-Scale Caching Infrastructure, Container Images, Gpu Device Access, Nvidia Container Toolkit, Cdi

Shield AI

Shield AI

San Diego, CA
Manager, Software Engineering
$170k+/yrOn-site7+ YOEEngineering Management

Leads a hands-on Shared Services Engineering team building and operating reusable services, SDKs, APIs, and customer-facing systems. Requires 7+ years of software engineering experience, engineering management experience, and strong technical judgment across distributed and full-stack systems.

Crusoe

Crusoe

Denver, CO

Senior Manager, Commissioning
$160k+/yrOn-site10+ YOEEngineering Management

Leads commissioning programs across multiple data center projects, managing commissioning teams, third-party agents, and stakeholder coordination from pre-functional testing through turnover. Requires 10+ years of mission-critical commissioning experience, 5+ years of leadership, engineering knowledge, and a bachelor’s degree.

LegitScript

LegitScript

United States

Manager, Platform Engineering
$160k+/yrRemote8+ YOEEngineering Management

Leads a hands-on platform engineering team responsible for AWS infrastructure, Kubernetes deployment paths, developer self-service, CI/CD governance, reliability, and audit readiness. The role requires deep infrastructure experience, Terraform expertise, production Kubernetes operations, and people leadership.

BuildOps

BuildOps

San Francisco, CA
Senior Engineering Manager, Financial Platform
$172k+/yrHybrid10+ YOEEngineering Management

Leads the architecture and development of BuildOps’ integration and data platform, driving technical strategy, reliability, APIs, databases, and engineering standards across multiple teams. Requires 10+ years of software engineering experience and deep expertise in TypeScript, Node.js, PostgreSQL, cloud infrastructure, and distributed systems.

Crusoe

Crusoe

Shakopee, MN

Senior Manager, Data Center Facility Operations
$175k+/yrOn-site5+ YOEEngineering Management

Leads 24/7 critical facility operations for a 20 MW AI data center expanding to 40 MW, overseeing electrical, mechanical, safety, maintenance, staffing, and operational performance. Requires at least five years of data center operations management experience and expertise in infrastructure reliability and team scaling.