Skip to content
DatadogDatadog

AI Research Engineer - Datadog AI Research

Builds data pipelines, distributed training infrastructure, simulation environments, and evaluation systems for multimodal world models and reinforcement-learning agents. The role partners with research scientists to scale experiments and turn prototypes into reliable observability and security products.

About the job

Responsibilities

  • Build and operate multimodal data pipelines, training and evaluation infrastructure, benchmarks, and internal tooling.
  • Implement models, run experiments at scale, and profile for reliability, performance, and cost.
  • Build simulation environments and replay infrastructure for agent training and evaluation.
  • Orchestrate distributed training and distributed reinforcement learning with Ray, including scheduling, scaling, and failure recovery.
  • Establish rigorous automated benchmarks and regression tests for world model predictions, agent performance, and simulation fidelity.
  • Collaborate with Research Scientists, Product, and Engineering to integrate capabilities into products and harden prototypes into reliable services.
  • Contribute to research publications at top-tier conferences such as NeurIPS, ICLR, and ICML.
  • Produce high-quality code, documentation, and open-source artifacts.

Requirements

  • Depth in distributed computing, reinforcement-learning infrastructure, and machine-learning systems for training and inference at scale.
  • Experience with Ray, Slurm, or similar frameworks is a plus.
  • Proficiency in Python and familiarity with a systems language such as Rust, C++, or Go.
  • Comfort with modern cloud and data infrastructure.
  • Practical experience implementing and operating ML training and inference systems such as PyTorch or JAX.
  • Experience with containerization, orchestration, and GPU acceleration.
  • Experience with large-scale model training and fine-tuning, including frameworks such as Megatron-LM, DeepSpeed, SkyRL, VeRL, or TorchTitan.
  • Familiarity with SFT, RLVR, RLHF, quantization, and speculative decoding.
  • Ability to explain design and performance trade-offs clearly to technical and non-technical audiences.
  • Experience supporting or contributing to research publications.

Nice-to-haves

  • Strong software engineering skills in observability, SRE, or security.
  • Experience bridging research prototypes and real-world product applications, especially with large foundation models, world models, or RL-trained agents.
  • Hands-on GPU programming and optimization, including CUDA.
  • Experience writing production data pipelines and applications.
  • Experience building simulation or sandbox environments for agent training.

Benefits

  • Competitive global benefits.
  • New-hire stock equity (RSUs) and employee stock purchase plan (ESPP).
  • Collaboration with colleagues across Datadog offices in New York City and Paris.
  • Opportunities to attend and present at conferences and meetups.
  • Intra-departmental mentor and buddy program.
  • Inclusive company culture and access to Community Guilds.

Skills

Python, Ray, Slurm, Rust, C++, Go, PyTorch, JAX, Megatron-Lm, Deepspeed, CUDA, Kubernetes, RLHF, Quantization, Speculative Decoding

Improbable

Improbable

Remote

AI Researcher
No salary listedRemoteAI Research

Conduct applied research on AI agents, designing experiments and evaluation systems to improve reliability, context retention, and multi-step task completion. The role requires strong AI/ML research, engineering, experimental design, and communication skills.

AI Digest

AI Digest

Remote

Research Scientist - Member of Technical Staff
$150k+/yrRemoteAI Research

Conduct research on long-horizon, multi-agent AI behavior by designing agent environments, analyzing large-scale data, and running experiments. The role requires strong research judgment, rapid execution, independence, and familiarity with current AI developments.

AI Digest

AI Digest

Remote

Engineer - Member of Technical Staff
$150k+/yrRemoteAI Research

Build, optimize, and evaluate long-running and multi-agent AI systems, along with tools for monitoring and analyzing their real-world behavior. The role requires software engineering experience with coding agents, strong independence, and familiarity with current AI developments.

Datadog

Datadog

Paris, France

Senior Applied Scientist - AI Platform
No salary listedHybrid6+ YOEAI Research

The applied scientist will lead GenSim’s methodology for creating realistic, production-grade simulated environments and high-quality post-training data for Datadog agents. The role requires deep LLM and agent experience, evaluation expertise, Python, distributed systems, and the ability to set technical direction.

Vanta

Vanta

Remote

Senior Product Builder, Organizational Intelligence
$176k+/yrRemote5+ YOEAI Research

Build Vanta’s organizational intelligence layer by shipping prototypes, internal tools, and AI agent workflows that make cross-source data useful to EPD, GTM, and other teams. The role requires recent hands-on LLM product work, independent problem scoping, and strong judgment around AI quality, reliability, cost, and latency.