Skip to content
DatabricksDatabricks

Principal Research Scientist – Scaling

Leads research team advancing LLM scaling, post-training, RL, and inference efficiency. Drives innovations in optimization, distributed systems, and production integration using Python/PyTorch, with deep expertise in large-scale ML.

About the job

Responsibilities

  • Lead and grow a multidisciplinary research team focused on foundational and applied AI problems, with emphasis on LLM scaling, efficiency, and systems performance.
  • Define the scaling research roadmap aligned with strategic objectives, prioritizing foundation model efficiency and large-scale training/inference.
  • Drive algorithmic innovations for large-scale neural network training/inference, including novel optimizers, low-precision techniques, and model adaptation methods.
  • Optimize end-to-end ML systems for distributed training/RL, memory/compute efficiency via collaboration with systems/platform teams.
  • Partner with product/engineering to translate research into customer-impacting capabilities.
  • Foster scientific excellence, reproducible experimentation, and knowledge sharing.
  • Represent research externally via publications, talks, and collaborations.
  • Mentor and develop research scientists/engineers.

What You Will Do

  • Define/lead research programs on foundation model efficiency (optimizer design, low-precision training/inference, scalable architectures, efficient adaptation).
  • Oversee large-scale experiments, benchmarking, and trade-off evaluation (quality, latency, throughput, cost).
  • Work hands-on with Python/PyTorch for research implementation, prototyping, and production integration.
  • Collaborate on distributed training, parallelism, memory management, hardware utilization.
  • Establish metrics/evaluation protocols for scaling research (training efficiency, inference cost, energy usage).
  • Champion responsible deployment ensuring model reliability/safety.

Requirements

  • Proven leadership of research teams developing novel foundation model efficiency techniques with industry impact.
  • Deep expertise in generative AI, LLMs, distributed ML systems, model optimization, or responsible AI, emphasizing scaling/efficiency.
  • Hands-on leadership with strong Python/PyTorch programming skills.
  • Ability to translate research into scalable product capabilities.
  • Excellent communication, leadership, stakeholder management skills.

Nice to Have

  • Experience at systems/ML intersection (distributed training frameworks, compiler/kernel optimization, memory/compute-efficient design).
  • Strong network in large-scale ML with conference service/collaborations.
  • Record of research impact (top ML/systems publications, open-source contributions, deployed systems).

Skills

PyTorch, Python, LLMs, Distributed Training, Model Optimization, Low-Precision Training, Rl, Neural Networks, Foundation Models, Scalable Architectures

Order.co

Order.co

Boston, MA
Principal Applied AI Architect
No salary listedRemote14+ YOEAI Research

Sets company-wide architecture and strategy for data and applied AI, connecting governed data foundations to production intelligence and measurable business outcomes. The role requires 14+ years of experience, strong production engineering judgment, executive partnership, and hands-on delivery.

Shield AI

Shield AI

San Mateo, CA

Senior Staff Software Engineer, Autonomy Capabilities
$281k+/yrOn-site10+ YOEAI Research

Leads design, implementation, integration, and field validation of tactical autonomy software for unmanned systems and multi-agent missions. Requires extensive autonomy or robotics experience, strong C++ and Python skills, and the ability to obtain a SECRET clearance.

Ambral

Ambral

New York, NY
Head of Research
$250k+/yrOn-site8+ YOEAI Research

Leads the research agenda and hands-on development of replayable enterprise environments, agent evaluations, and post-training systems. The role requires deep AI research experience, a PhD or equivalent track record, and the ability to translate open-ended questions into production systems.

Shield AI

Shield AI

Washington, DC
Senior Staff Engineer, Autonomy Capabilities – Maritime
$221k+/yrOn-site10+ YOEAI Research

Leads technical direction and develops maritime autonomy capabilities for unmanned surface and underwater vehicles, including motion planning, localization, safe behaviors, and heterogeneous multi-agent collaboration. Requires deep robotics and unmanned-systems experience, strong C++/Python skills, and senior technical leadership.

Deepgram

Deepgram

San Francisco, CA
Director of Research, Text to Speech
$213k+/yrRemote8+ YOEAI Research

Leads Deepgram’s end-to-end TTS research program, setting technical direction, training and evaluating large-scale speech-generation models, and turning breakthroughs into production systems. The role combines hands-on technical leadership with building and developing a high-performing research organization.