Skip to content
DatabricksDatabricksSan Francisco, CA

Staff Backend Software Engineer- (AI Platform)

Staff Backend Engineer building and leading the architecture of Databricks Model Serving platform for high-throughput, low-latency AI/ML inference on CPU and GPU. Requires 10+ years large-scale distributed systems experience and deep inference infrastructure expertise.

192k – 260k/yr
On-site10+ YOEML Engineering

About the role

Responsibilities

  • Design and implement core systems and APIs that power Databricks Model Serving, ensuring scalability, reliability, and operational excellence.
  • Partner with product and engineering leadership to define the technical roadmap and long-term architecture for serving workloads.
  • Drive architectural decisions and trade-offs to optimize performance, throughput, autoscaling, and operational efficiency for CPU and GPU serving workloads.
  • Contribute directly to key components across the serving infrastructure — from model container builds and deployment workflows to runtime systems like routing, caching, observability, and intelligent autoscaling — ensuring smooth and efficient operations at scale.
  • Collaborate cross-functionally with product, platform, and research teams to translate customer needs into reliable and performant systems.
  • Lead technical initiatives that improve latency, availability, and cost-effectiveness across both customer-facing and foundational serving layers.
  • Establish best practices for code quality, testing, and operational readiness, and mentor other engineers through design reviews and technical guidance.
  • Represent the team in cross-organizational technical discussions and influence Databricks’ broader AI platform strategy.

Requirements

  • 10+ years of experience building and operating large-scale distributed systems.
  • Deep expertise in model serving, inference systems, and related infrastructure (e.g., routing, scheduling, autoscaling, and observability).
  • Strong foundation in algorithms, data structures, and system design as applied to large-scale, low-latency serving systems.
  • Proven ability to deliver technically complex, high-impact initiatives that create measurable customer or business value.
  • Experience leading architecture for large-scale, performance-sensitive CPU/GPU inference systems.
  • Strong communication skills and ability to collaborate across teams in fast-moving environments.
  • Strategic and product-oriented mindset with the ability to align technical execution with long-term vision.
  • Passion for mentoring, growing engineers, and fostering technical excellence.

Skills

Distributed Systemsmodel servinginference systemsroutingschedulingautoscalingObservabilitySystem DesigncpuGPUAlgorithmsData Structures

Similar roles

ML Engineering jobs
Databricks

Staff Software Engineer

DatabricksNew York, NY

Staff ML Engineer building CustomerLake, Databricks' Customer Data Platform for enterprise ML/AI personalization, recommendations, churn, and LTV modeling. Requires 10+ years shipping production ML/LLM systems with strong product mindset in 0-to-1 environments.

192k – 260k/yr
On-site10+ YOEML Engineering
Databricks

Staff Software Engineer, Model Serving

DatabricksSan Francisco, CA

Designs and builds scalable, low-latency model serving infrastructure for AI/ML models across CPU/GPU workloads. Requires 10+ years in large-scale distributed systems and deep expertise in inference systems, architecture, and cross-team collaboration.

192k – 260k/yr
On-site10+ YOEML Engineering
Databricks

Staff Software Engineer, Foundational Model Serving

DatabricksSan Francisco, CA

Designs and builds scalable, low-latency systems for serving frontier AI models on GPUs. Requires 10+ years in large-scale distributed systems, strong system design skills, and leadership in operational excellence; no prior AI experience needed.

192k – 260k/yr
On-site10+ YOEML Engineering
Databricks

Staff Software Engineer - GenAI Performance and Kernel

DatabricksSan Francisco, CA

Designs, implements, and optimizes high-performance GPU kernels for GenAI inference stack. Leads performance improvements, mentors engineers, and collaborates with ML and systems teams. Requires deep kernel programming and GPU architecture expertise.

191k – 233k/yr
On-siteML Engineering
Databricks

Staff Software Engineer - GenAI inference

DatabricksSan Francisco, CA

Leads architecture, development, and optimization of GenAI inference engine for high-throughput, low-latency LLM serving. Requires 6+ years in performance-critical systems, deep ML inference expertise, CUDA/GPU programming, and distributed systems.

191k – 233k/yr
On-site6+ YOEML Engineering