# Staff Backend Software Engineer- (AI Platform)

**Company:** [Databricks](https://hotfix.jobs/companies/databricks)
**Location:** San Francisco, CA
**Role:** ML Engineering
**Salary:** $166k – $225k/yr
**Experience:** 7+ years
**Skills:** Distributed Systems, model serving, inference systems, routing, scheduling, autoscaling, Observability, System Design, cpu, GPU, Algorithms, Data Structures
**Posted:** 2026-07-17

> Build and optimize core infrastructure for Databricks Model Serving, focusing on high-throughput, low-latency inference for CPU/GPU workloads. Requires 5+ years in large-scale distributed systems and experience with model serving, inference, routing, autoscaling and observability.

## Job Description

## Impact
- Design and implement core systems and APIs that power Databricks Model Serving, ensuring scalability, reliability, and operational excellence.
- Drive architectural decisions and trade-offs to optimize performance, throughput, autoscaling, and operational efficiency for CPU and GPU serving workloads.
- Contribute directly to key components across the serving infrastructure — from model container builds and deployment workflows to runtime systems like routing, caching, observability, and intelligent autoscaling — ensuring smooth and efficient operations at scale.
- Collaborate cross-functionally with product, platform, and research teams to translate customer needs into reliable and performant systems.
- Lead technical initiatives that improve latency, availability, and cost-effectiveness across both customer-facing and foundational serving layers.
- Establish best practices for code quality, testing, and operational readiness, and mentor other engineers through design reviews and technical guidance.

## Requirements
- 5+ years of experience building and operating large-scale distributed systems.
- Experience in model serving, inference systems, or related infrastructure (e.g., routing, scheduling, autoscaling, and observability).
- Strong foundation in algorithms, data structures, and system design as applied to large-scale, low-latency serving systems.
- Proven ability to deliver technically complex, high-impact initiatives that create measurable customer or business value.
- Experience building architecture for large-scale, performance-sensitive CPU/GPU inference systems.
- Strong communication skills and ability to collaborate across teams in fast-moving environments.
- Customer-focused mindset with the ability to align implementation details with product goals.
- Passion for mentoring, growing engineers, and fostering technical excellence.

## Similar roles

- [Senior Staff AI & Agentic Systems Engineer](https://hotfix.jobs/jobs/5aca3ad2-37ed-432f-acc7-d7916c850288) - Mozilla - Remote - $166k – $260k/yr
- [Staff Software Engineer, AI Search](https://hotfix.jobs/jobs/0fb6896f-890f-465e-8866-968f76175ae7) - Databricks - Mountain View, CA - $166k – $225k/yr
- [Mid/Senior/Staff Software Engineer, Agents](https://hotfix.jobs/jobs/77110eb4-a2f2-4725-98fc-709a6b518954) - Harvey - San Francisco, CA - $165k – $312k/yr
- [Retirement AI Staff Engineer](https://hotfix.jobs/jobs/307cb3aa-0135-4f07-9cbd-f1109c32fd31) - Gusto - Denver, CO - $164k – $247k/yr
- [Senior / Staff Software Engineer](https://hotfix.jobs/jobs/f23e3c2b-6a0e-4f62-a932-7681b5efa73e) - Clear Street - Remote - $170k – $240k/yr

**Apply:** https://hotfix.jobs/jobs/0371944a-14f8-489e-9041-b504f082128c
**Canonical:** https://hotfix.jobs/jobs/0371944a-14f8-489e-9041-b504f082128c