Skip to content
DatabricksDatabricks

Sr. Engineering Manager, AI Runtime

Lead engineering team building Databricks AI Runtime for large-scale GPU model training and fine-tuning. Own product experience, distributed training infrastructure, reliability, and roadmap while collaborating across product, research, and platform teams. Requires 8+ years software engineering and 3+ years management experience with deep GPU training expertise.

About the job

Impact

  • Lead, mentor, and grow a high-performing engineering team responsible for the Custom Training product and its foundational infrastructure, including distributed training orchestration, cluster lifecycle, fault tolerance, and training efficiency.
  • Define and own the product and technical roadmap for AIR, balancing customer experience, functionality, and foundational investments.
  • Collaborate closely with product, research, platform, infrastructure teams, and customers to drive end-to-end delivery, from ideation and prioritization to launch and operation.
  • Drive architectural decisions and product design for managed GPU training at scale.
  • Advocate for customer needs through direct engagement, ensuring engineering decisions translate to clear product impact.
  • Build observability and reliability practices for long-running, multi-node training jobs, including checkpoint strategies, failure recovery, and operational runbooks.
  • Partner with recruiting to attract, hire, and develop top-tier engineering talent.

Requirements

  • 8+ years of software engineering experience, with 3+ years in engineering management.
  • Track record building and operating managed GPU training infrastructure at scale (100s/1000s GPUs).
  • Deep familiarity with distributed training frameworks (PyTorch, DeepSpeed, Composer, Megatron-LM) and parallelism strategies (FSDP, tensor/pipeline parallelism).
  • Experience with training resilience patterns: checkpointing, elastic training, and automated failure recovery for long-running jobs.
  • Understanding of GPU performance fundamentals including NCCL, interconnect topologies, and memory optimization.
  • Experience building platform products with clear SLAs where you've owned the customer experience, not just the backend.
  • Strong cross-functional leadership across platform, product, and research teams, with the ability to lead through ambiguity and deliver complex projects.
  • Excellent collaboration and communication skills across engineering, product, and research organizations.
  • BS/MS in Computer Science, Electrical Engineering, or related technical field.

Skills

PyTorch, Deepspeed, Megatron-Lm, Fsdp, Nccl, Gpu Training, Distributed Training, Checkpointing, Elastic Training, Tensor Parallelism, Pipeline Parallelism

Databricks

Databricks

Mountain View, CA

Senior Engineering Manager - Enzyme
$229k+/yrOn-site5+ YOEEngineering Management

Leads the engineering team responsible for next-generation Materialized View capabilities within the Lakeflow data infrastructure platform. The role requires engineering management experience plus deep expertise in data infrastructure, distributed systems, database internals, and team development.

Skydio

Skydio

San Mateo, CA

Senior Manager, Middleware
$230k+/yrHybrid8+ YOEEngineering Management

Leads and develops a software engineering team building middleware and on-device services for autonomous drone platforms. The role combines people management with hands-on Python or C++ development, systems architecture, reliability improvements, and cross-functional technical leadership.

Vercel

Vercel

San Francisco, CA
Member of the Technical Staff - Next.js
$230k+/yrHybrid8+ YOEEngineering Management

Hands-on engineering team lead building and evolving Next.js while coaching a small team and setting technical direction. Requires 8+ years of software engineering experience, strong React or framework architecture expertise, and continued production coding experience.

Crusoe

Crusoe

Bellevue, WA
Senior Manager, Network & Cloud Architecture
$230k+/yrOn-site10+ YOEEngineering Management

Leads the deployment organization responsible for bringing global HPC and GPU infrastructure online, owning people, delivery, standards, vendors, and executive-level program risk. Requires 10+ years in network engineering or data center deployment and substantial engineering management experience.

Genius AI

Genius AI

San Francisco, CA

Senior Engineering Manager, Infrastructure
$230k+/yrOn-site8+ YOEEngineering Management

Leads Compute, Data, and Developer Experience teams, owning infrastructure strategy, reliability, developer enablement, vendor spend, and team growth. The role requires 8+ years of infrastructure or platform experience, 4+ years of engineering leadership, and deep expertise in AWS, Kubernetes, Terraform, ArgoCD, and observability.