Skip to content
DatabricksDatabricksSan Francisco, CA

Staff Software Engineer

Build and optimize LLM inference infrastructure at enterprise scale for partner and self-hosted frontier models. Requires 8+ years backend/infrastructure engineering experience with distributed systems, real-time serving, and ML/GPU orchestration.

190k – 265k/yr
On-site8+ YOEML Engineering

About the role

Impact

  • Build LLM infrastructure powering large-scale inference workloads for customers through partner models (OpenAI, Anthropic, Gemini) and self-hosted models (Qwen, GPT-OSS, Llama)
  • Improve reliability, latency, and efficiency of distributed AI workloads
  • Collaborate with platform, infra, and ML teams to deliver seamless end-to-end experiences
  • Shape how developers and data scientists build and interact with AI on Databricks

Requirements

  • 8+ years of experience in backend or infrastructure engineering
  • Experience with distributed systems, scalable APIs, or cloud-native infrastructure
  • Experience with real-time serving, ML infrastructure, or GPU orchestration
  • Familiarity with service-oriented architecture, deployment pipelines, and system observability

Nice-to-Haves

  • Exposure to platforms like SageMaker, Vertex AI, or Azure ML
  • Contributions to OSS projects like MLflow, PyTorch, Ray, vLLM, SGLang
  • Built developer platforms or internal tools supporting AI workflows

Skills

Distributed Systemsscalable apiscloud native infrastructurereal-time servingML Infrastructuregpu orchestrationservice-oriented architecturedeployment pipelinessystem observabilitySageMakervertex aiazure mlMLflowPyTorchRay

Similar roles

ML Engineering jobs
Shield AI

Staff Engineer, AI Platform & Architecture

Shield AIUnited States

Staff Engineer responsible for designing enterprise AI platform architecture, reusable components, responsible AI controls, observability, and governance patterns. Requires deep expertise in generative AI, RAG, agentic workflows, and influencing cross-functional teams without direct authority.

190k – 290k/yr
Remote7+ YOEML Engineering
Databricks

Staff Software Engineer - AI Research Infrastructure

DatabricksNew York, NY

Founding member of a new team building foundational evaluation infrastructure and flywheels for Databricks' AI/Genie Agents. Design scalable tooling for benchmarking, regression detection, and quality measurement that drives continuous agent improvement across research, training, and production.

190k – 270k/yr
On-site6+ YOEML Engineering
Databricks

Staff Software Engineer, Foundation Model API

DatabricksSan Francisco, CA

Build and shape the Foundation Model API serving layer for large-scale LLM inference (partner and self-hosted models) at Databricks. Requires 8+ years backend/infra engineering experience with distributed systems, ML infrastructure, and a strong product ownership mindset.

190k – 265k/yr
On-site8+ YOEML Engineering
Databricks

Staff Software Engineer, AI Runtime

DatabricksMountain View, CA +1

Staff Software Engineer building and scaling Databricks' managed large-scale GPU training platform (AIR). Focus on distributed training performance, scheduling, fault tolerance, and developer experience for thousands of accelerators.

190k – 265k/yr
On-site10+ YOEML Engineering
Databricks

Staff Machine Learning Engineer

DatabricksSan Francisco, CA +1

Develops and deploys state-of-the-art GenAI models and systems for Databricks products like Assistant and Genie. Requires 2-8 years ML engineering experience, proficiency in Python/PyTorch/TensorFlow, and expertise in LLMs.

190k – 285k/yr
HybridML Engineering