Latest ML Engineering jobs at Databricks
Job results
Leads the development of ML- and NLP-powered search relevance systems, including query understanding, ranking, retrieval, and evaluation. Requires 10+ years of search relevance experience and a bachelor’s degree, with advanced study preferred.
Develop and productionize generative AI applications for customers while advising stakeholders and influencing product direction. The role requires extensive industry data science experience, production ML deployment expertise, graduate-level quantitative training or equivalent experience, and native Japanese with professional English.
Build and productionize generative AI applications for U.S. federal customers, advise clients, and influence product direction. The role requires extensive data science and machine learning deployment experience, a graduate quantitative degree or equivalent experience, and U.S. security clearance eligibility.
Senior software engineer developing ML-based search relevance and discovery systems, including query understanding, ranking, retrieval, and evaluation pipelines. The role requires 5+ years of search relevance experience and expertise in NLP, LLMs, or related discovery technologies.
Leads the development of machine-learning search relevance systems, including query understanding, ranking, retrieval, and evaluation pipelines. The role requires 10+ years of search relevance experience and expertise in NLP, LLMs, or related discovery technologies.
Develop and deploy scalable machine learning and AI systems for user-facing products, including forecasting and AutoML capabilities. The role requires strong production ML engineering, modeling, software engineering, statistics, and systems knowledge.
Leads the development of machine-learning search relevance systems, including query understanding, ranking, retrieval, and evaluation at scale. Requires 10+ years of search relevance experience and expertise in ML, NLP, or related discovery technologies.
The Senior ML and AI Technical Solutions Engineer troubleshoots and optimizes production data, machine learning, and generative AI workloads on Databricks. The role requires 8+ years of production experience with ML/AI systems, distributed computing, cloud platforms, and programming in Python, Scala, and Java.
This senior applied ML role builds and deploys optimization-driven machine learning systems for serverless infrastructure, spanning cluster management through query compilation. It requires production ML experience, cloud and distributed-systems knowledge, strong programming skills, and a master's degree in a related computational field.
Staff engineer driving technical vision and architecture for Spark Structured Streaming. Build core capabilities like advanced state management and operators; improve latency, throughput, and cost. Requires 8+ years in big-data, Spark, or database systems plus passion for distributed systems.
Build and ship AI-powered analytics agents, LLM-driven workflows, and intelligence features for Databricks' internal GTM platform (Customer Zero). Requires 7+ years software/AI engineering experience, strong Python/SQL, hands-on LLM/RAG/agents experience, and fluency with AI coding tools like Claude.
Build and optimize LLM inference infrastructure at enterprise scale for partner and self-hosted frontier models. Requires 8+ years backend/infrastructure engineering experience with distributed systems, real-time serving, and ML/GPU orchestration.
Founding member of a new team building foundational evaluation infrastructure and flywheels for Databricks' AI/Genie Agents. Design scalable tooling for benchmarking, regression detection, and quality measurement that drives continuous agent improvement across research, training, and production.
Build and shape the Foundation Model API serving layer for large-scale LLM inference (partner and self-hosted models) at Databricks. Requires 8+ years backend/infra engineering experience with distributed systems, ML infrastructure, and a strong product ownership mindset.
Staff Software Engineer owning architecture and delivery of LLM-powered agentic workflows for marketing content creation, publishing, and reliability at scale. Requires 12+ years experience building production LLM systems, human-in-the-loop designs, stakeholder collaboration, and mentoring.
Staff ML Engineer building CustomerLake, Databricks' Customer Data Platform for enterprise ML/AI personalization, recommendations, churn, and LTV modeling. Requires 10+ years shipping production ML/LLM systems with strong product mindset in 0-to-1 environments.
Staff Software Engineer building and scaling Databricks' managed large-scale GPU training platform (AIR). Focus on distributed training performance, scheduling, fault tolerance, and developer experience for thousands of accelerators.
Senior Software Engineer building and scaling Databricks' managed GPU training platform (AI Runtime) for large-scale distributed AI model training. Requires 5+ years in distributed systems and hands-on experience with GPU training frameworks.
As a Staff Engineer for Search, you will build and scale Databricks' next-generation Search product, driving the design and evolution of a highly-performant, cost-efficient, and developer-friendly Search stack. You will also define the long-term vision, mentor senior engineers, and lead strategic efforts.
Builds retrieval stack and search subagents for Databricks AI agents, handling query understanding, hybrid retrieval across structured/unstructured enterprise data, and evaluation. Requires 10+ years in production IR/RAG systems and agentic workflows.
Develops and deploys state-of-the-art GenAI models and systems for Databricks products like Assistant and Genie. Requires 2-8 years ML engineering experience, proficiency in Python/PyTorch/TensorFlow, and expertise in LLMs.
Designs, implements, and optimizes high-performance GPU kernels for GenAI inference stack. Leads performance improvements, mentors engineers, and collaborates with ML and systems teams. Requires deep kernel programming and GPU architecture expertise.
Leads architecture, development, and optimization of GenAI inference engine for high-throughput, low-latency LLM serving. Requires 6+ years in performance-critical systems, deep ML inference expertise, CUDA/GPU programming, and distributed systems.
Designs and builds scalable, low-latency model serving infrastructure for AI/ML models across CPU/GPU workloads. Requires 10+ years in large-scale distributed systems and deep expertise in inference systems, architecture, and cross-team collaboration.
Designs and builds scalable, low-latency systems for serving frontier AI models on GPUs. Requires 10+ years in large-scale distributed systems, strong system design skills, and leadership in operational excellence; no prior AI experience needed.
Designs and builds scalable infrastructure for high-throughput, low-latency AI/ML model serving on CPU/GPU. Requires 5+ years in distributed systems, inference expertise, and strong system design skills.