Staff Backend Software Engineer- (AI Platform)
Builds infrastructure powering Databricks' AI platform including MLflow, AI Gateway, and model serving. Requires 5+ years backend experience with Scala, Go, or Python, and expertise in distributed systems and scalable APIs.
About the job
Impact/Responsibilities
- Build infrastructure that powers flagship offerings like MLflow, AI Gateway, Databricks Apps, Agent Framework, Agent Bricks, and Foundation Model APIs.
- Improve reliability, latency, and efficiency of distributed AI workloads.
- Collaborate with platform, infra, and ML teams to deliver seamless end-to-end experiences.
- Shape how developers and data scientists build and interact with AI on Databricks.
Requirements
- 5+ years of experience in backend or infrastructure engineering.
- Strong programming skills in Scala, Go, or Python.
- Experience with distributed systems, scalable APIs, or cloud-native infrastructure.
- Familiarity with service-oriented architecture, deployment pipelines, and system observability.
- Strong product and ownership mindset.
Nice-to-Haves
- Experience with real-time serving, ML infrastructure, or GPU orchestration.
- Exposure to platforms like SageMaker, Vertex AI, or Azure ML.
- Contributions to OSS projects like MLflow, PyTorch, or Ray.
- Built developer platforms or internal tools supporting AI workflows.
Skills
Scala, Go, Python, Distributed Systems, Scalable Apis, Cloud-Native Infrastructure, Service-Oriented Architecture, MLflow, ML Infrastructure, Gpu Orchestration
Similar jobs
Backend Engineering jobsTechnical leader for the Metadata team, designing distributed cloud subsystems and leading complex initiatives across discovery, catalog, lineage, and run history services. Requires 8+ years of software engineering experience, backend or systems expertise, cloud infrastructure experience, and strong architecture, reliability, and mentoring skills.
Technical leader for the Metadata team, owning architecture and delivery of distributed backend services for discovery, catalog, lineage, and run history. The role requires 8+ years of software engineering experience, strong cloud and systems expertise, and a record of leading multi-engineer initiatives.
Designs and operates Internet-scale scanning, DNS, attribution, and data pipelines, with deep ownership of distributed backend systems and production reliability. Requires 10+ years of software engineering experience, strong Go expertise, cloud and streaming infrastructure knowledge, and the ability to mentor engineers.
Staff backend engineer responsible for evolving Grafana into a scalable, multi-tenant observability application platform. The role requires production operations experience, distributed-systems expertise, strong communication, and familiarity with Go or willingness to learn it.
Leads architecture and development of mission-critical backend systems for spacecraft command, telemetry, mission planning, and operations. Requires 8+ years of software development experience, distributed-systems and cloud-native expertise, and technical leadership.