Skip to content
DatabricksDatabricks

Sr. Staff AI Research TLM - AI Systems

Lead a world-class research team advancing LLM scaling, post-training, RL, and inference efficiency at Databricks AI. Drive research roadmap and translate breakthroughs into production systems while collaborating closely with engineering and product teams.

About the job

Responsibilities

  • Lead and grow a multidisciplinary research team focused on foundational and applied AI problems, with emphasis on LLM scaling, efficiency, and systems performance.
  • Define the scaling research roadmap in alignment with Databricks’ strategic objectives, prioritizing advances in foundation model efficiency and large-scale training and inference.
  • Drive algorithmic innovations for large-scale neural network training and inference, including novel optimizers, low-precision techniques, and model adaptation methods, and guide rigorous empirical validation.
  • Optimize end-to-end ML systems for distributed training and RL, memory efficiency, and compute efficiency through collaboration with core systems and platform teams.
  • Partner with product and engineering to translate research breakthroughs into customer-impacting capabilities in the Databricks AI platform.
  • Foster a culture of scientific excellence, reproducible experimentation, and internal knowledge sharing.
  • Represent Databricks AI research externally through top-tier publications, conference talks, and collaborations with academia and open-source community.
  • Mentor and develop talent, providing technical guidance and career development support.

Key Requirements

  • Proven ability to lead a research team to develop novel techniques for foundation model efficiency, with strong track record of industry impact.
  • Deep expertise in at least one of: generative AI, LLMs, distributed ML systems, model optimization, or responsible AI, with emphasis on scaling and efficiency for large-scale neural networks.
  • Strong programming skills and demonstrated ability to write high-quality, efficient code in Python and PyTorch for research implementation and experimentation.
  • Demonstrated ability to translate research innovation into scalable product capabilities in partnership with product and engineering teams.
  • Excellent communication, leadership, and stakeholder management skills.

Nice-to-Haves

  • Prior work at the intersection of systems and ML, such as distributed training frameworks, compiler/kernel optimization for deep learning workloads, or memory-/compute-efficient model design.
  • Strong industry and academic network in large-scale ML, with collaborations or service at top conferences in ML and systems.
  • Strong record of research impact—first-author publications at top ML/systems conferences (e.g., ICLR, ICML, NeurIPS, MLSys), influential open-source contributions, or widely used deployed systems.

Skills

Python, PyTorch, LLMs, Distributed Ml Systems, Model Optimization, Low-Precision Training/Inference, Rl (Reinforcement Learning), Foundation Model Efficiency, Distributed Training Frameworks, Compiler/Kernel Optimization

Shield AI

Shield AI

San Mateo, CA

Senior Staff Software Engineer, Autonomy Capabilities
$281k+/yrOn-site10+ YOEAI Research

Leads design, implementation, integration, and field validation of tactical autonomy software for unmanned systems and multi-agent missions. Requires extensive autonomy or robotics experience, strong C++ and Python skills, and the ability to obtain a SECRET clearance.

Shield AI

Shield AI

Washington, DC
Senior Staff Engineer, Autonomy Capabilities – Maritime
$221k+/yrOn-site10+ YOEAI Research

Leads technical direction and develops maritime autonomy capabilities for unmanned surface and underwater vehicles, including motion planning, localization, safe behaviors, and heterogeneous multi-agent collaboration. Requires deep robotics and unmanned-systems experience, strong C++/Python skills, and senior technical leadership.

Upstart

Upstart

United States

Staff Machine Learning Model Risk Specialist
$140k+/yrRemote7+ YOEAI Research

Evaluates model and Generative AI risks across Upstart Bank’s model inventory, conducting risk assessments, monitoring reviews, quantitative analyses, and governance activities. Requires a quantitative master’s degree, 4+ years of relevant experience, and coding skills in Python, R, or similar languages.

Anthropic

Anthropic

San Francisco, CA

Staff+ Researcher, Cybersecurity Products
$405k+/yrHybrid7+ YOEAI Research

Research and evaluate frontier AI capabilities for cybersecurity, rapidly prototyping tools, designing rigorous benchmarks, and helping operationalize reliable capabilities into products. Requires deep security expertise, strong technical communication, and at least seven years of relevant experience.

Order.co

Order.co

United States

Staff Applied AI Scientist
No salary listedRemote10+ YOEAI Research

Own the architecture, delivery, evaluation, and production operations of AI capabilities embedded in procurement and finance workflows. The role requires 10+ years in applied AI or machine learning, deep LLM and agent expertise, and experience delivering measurable production outcomes.