Skip to content
DatabricksDatabricks

Senior Software Engineer, AI Runtime

Senior Software Engineer building and scaling Databricks' managed GPU training platform (AI Runtime) for large-scale distributed AI model training. Requires 5+ years in distributed systems and hands-on experience with GPU training frameworks.

About the job

Responsibilities

  • Drive the architecture and evolution of AIR's managed GPU training platform, delivering scalable, high-throughput, and resilient training across fleets that span thousands of accelerators.
  • Solve the hardest problems in large-scale training, including multi-node orchestration, distributed parallelism strategies, GPU scheduling and dynamic routing, high-throughput data loading, and checkpoint and restore for very long-running jobs.
  • Push GPU efficiency and training performance, raising utilization (such as model FLOPs utilization and end-to-end throughput) and lowering cost per training run across diverse model architectures and hardware generations.
  • Build the resilience and observability foundations that keep multi-node jobs healthy, detecting and recovering from hardware and software failures with minimal disruption to customers.
  • Partner with product, research, and platform teams to shape the APIs, CLI, and developer experience that make it easy to launch, monitor, and debug production training jobs.
  • Lead end-to-end engineering efforts, from design through production rollout, holding a high bar for performance, correctness, and reliability.
  • Make direct, high-impact contributions to the core systems behind AIR, and help bring up support for the latest accelerators and new regions as the fleet grows.
  • Champion engineering excellence, mentor other engineers through design reviews and technical discussions, and contribute to Databricks' technical direction in AI training infrastructure.

Requirements

  • 5+ years of experience building and operating large-scale distributed systems, with experience in GPU training infrastructure, high-performance computing, or ML systems.
  • Experience with distributed training frameworks (such as PyTorch, FSDP, DeepSpeed, or Megatron) and the parallelism strategies (data, tensor, pipeline, and sequence parallelism) used to train large models.
  • Strong understanding of training resilience patterns, including checkpointing, failure detection, and automatic recovery for long-running, multi-node jobs.
  • Solid grasp of GPU performance fundamentals, including accelerator architecture, high-speed interconnects (such as NVLink and InfiniBand or RoCE), collective communication, and the bottlenecks that govern training throughput and utilization.
  • Experience building and operating managed, multi-tenant platform products in the cloud, with clear SLAs and SLOs for availability, performance, and reliability.
  • Strong foundation in algorithms, data structures, and system design as applied to performance-sensitive, large-scale distributed systems.
  • Proven ability to deliver technically complex, high-impact initiatives that create clear customer or business value.
  • Strong communication skills and the ability to collaborate across product, research, and infrastructure teams in a fast-moving environment.
  • Customer-focused mindset with the ability to align implementation details with product goals, and a passion for mentoring engineers and fostering technical excellence.
  • BS in Computer Science or a related field (MS or PhD preferred).

Nice-to-Haves

  • MS or PhD in Computer Science or a related field.

Skills

PyTorch, Fsdp, Deepspeed, Megatron, Gpu Scheduling, Distributed Training, Checkpointing, Nvlink, InfiniBand, Roce, Collective Communication, Multi-Node Orchestration, High-Performance Computing, Ml Systems, System Design

LangChain

LangChain

New York, NY

Lead Applied AI Engineer
$160k+/yrOn-site7+ YOEML Engineering

Leads hands-on development and deployment of production AI agents and workflows, sets technical standards, and mentors engineers. Requires 4+ years of software engineering experience, production LLM application experience, and strong Python or TypeScript skills.

LegitScript

LegitScript

United States

Senior Data Science Engineer
$160k+/yrRemote5+ YOEML Engineering

Owns the full lifecycle of data and ML solutions, from ingestion and feature-ready datasets through production deployment and business-impact measurement. The role combines data engineering, applied machine learning, MLOps, and generative AI to build risk detection capabilities.

Mercury

Mercury

San Francisco, CA
Senior Machine Learning Operations Engineer
$157k+/yrRemote5+ YOEML Engineering

Build and operate the platform that deploys, serves, observes, and retrains production machine-learning models for real-time fraud and financial-crime risk decisions. Requires 5+ years of ML engineering, backend, or MLOps experience, strong Python skills, and production model-serving expertise.

Parallel Systems

Parallel Systems

Los Angeles, CA

Senior Robotics Software Engineer, Perception
$163k+/yrOn-site5+ YOEML Engineering

Develop and deploy real-time perception and sensor-fusion software for autonomous battery-electric rail vehicles. The role requires strong robotics, geometry-based computer vision, C/C++ and Rust experience, plus hands-on work with multimodal sensors and production systems.

Lightning AI

Lightning AI

New York, NY
Senior Research Engineer, LLM Training & Post-Training
$165k+/yrHybrid5+ YOEML Engineering

Build and optimize large language model training and post-training pipelines, improving model quality, distributed performance, evaluation, and production readiness. The role requires deep PyTorch and transformer experience, strong distributed-systems and software-engineering skills, and expertise in modern LLM optimization techniques.