AI Research Scientist – Datadog AI Research (DAIR)
Conducts cutting-edge research in Generative AI, building foundation models and autonomous agents for cloud observability, SRE, and code repair. Requires PhD in ML or related field, publications at top conferences, and expertise in PyTorch/TensorFlow distributed training.
About the job
What You’ll Do
- Conduct cutting-edge research in Generative AI and Machine Learning to build specialized Foundation Models and AI Agents for observability, SRE, and code repair.
- Leverage large-scale distributed training infrastructure to pre-train and post-train state-of-the-art models on diverse telemetry data.
- Build simulated environments for on-policy agentic training and evaluation.
- Lead and contribute to research publications at top conferences (NeurIPS, ICLR, ICML) and open-source model artifacts and benchmarks.
- Collaborate with cross-functional teams to integrate AI capabilities into Datadog’s products.
- Stay at the forefront of LLMs, Foundation Models, and Generative AI research.
- Foster scientific rigor, innovation, and practical impact through reading groups and mentoring.
Who You Are
- PhD in Computer Science, Machine Learning, or related field with expertise in generative modeling, AI agents, reinforcement learning, or NLP (or equivalent experience).
- Extensive experience designing and implementing deep learning models and agents; strong background in distributed training (DeepSpeed, Megatron-LM) and ML libraries (PyTorch, TensorFlow).
- Proven track record of impactful research with publications at top venues (NeurIPS, ICLR, ICML, TMLR).
- Familiar with efficient training, post-training, fine-tuning, and inference for large foundation models.
- Excel at explaining complex models to technical and non-technical audiences.
- Strong interest in open-science and open-source contributions.
Bonus Points
- Ability to bridge research and real-world product applications, especially foundation models, generative AI agents, or domain-specific LLMs.
- Passion for customer impact, scalability, and responsible AI deployment.
- Experience writing production data pipelines and applications.
- Hands-on GPU programming and optimization, including CUDA.
Skills
PyTorch, TensorFlow, Deepspeed, Megatron-Lm, Generative AI, Foundation Models, Reinforcement Learning, AI Agents, Distributed Training, CUDA
Similar jobs
AI Research jobsResearch Engineer focused on designing benchmarks, evaluation systems, rubrics, and failure-analysis workflows for frontier language models. The role requires strong applied AI research and coding experience, with expertise in model evaluation, data quality, and backend systems.
Conduct research on long-horizon, multi-agent AI behavior by designing agent environments, analyzing large-scale data, and running experiments. The role requires strong research judgment, rapid execution, independence, and familiarity with current AI developments.
Build, optimize, and evaluate long-running and multi-agent AI systems, along with tools for monitoring and analyzing their real-world behavior. The role requires software engineering experience with coding agents, strong independence, and familiarity with current AI developments.
Research Scientist developing and evaluating health-focused AI models, large language models, and agentic systems for clinical applications. The role requires advanced research experience, strong coding skills, healthcare or clinical-data experience, and top-tier AI/ML publications.
Develops experimental AI techniques and prototypes for agentic marketing applications, with emphasis on image and video generation. The role requires strong backend or probabilistic systems expertise, quantitative thinking, creativity with LLM applications, and product intuition.