Staff Applied Scientist - Agentic Interfaces
As a Staff Applied Scientist, you will define and build measurement systems for AI agent interfaces at Datadog, focusing on evaluation strategy, metric definition, and dataset creation to improve agent performance on customer workflows.
About the job
What You’ll Do:
- Own the evaluation strategy for Datadog's AI agent integrations. Define the metrics — offline and online, quality and cost, single-turn and trajectory-level — that the team and the broader organization optimize against.
- Build the eval datasets, golden traces, and regression harnesses that catch quality changes before they hit customers, and make those assets reusable by every team contributing tools to the platform.
- Drive measurable improvements to retrieval relevance, tool-selection accuracy, and context efficiency, partnering closely with the AI engineers on the team who build the underlying platform.
- Run applied research on the open problems in agent–data interaction: tool selection under large catalogs, multi-turn agent evaluation, grounding and hallucination control on live telemetry, cost/quality tradeoffs at scale.
- Partner with the Bits SRE, Bits Assistant, and Bits Dev Agent teams so first-party agents benefit from the same measurement substrate as third-party integrations, and so learnings move freely in both directions.
- Provide technical leadership across the Agentic Interfaces team and the broader organization through design reviews, working groups, and mentorship, and represent the team externally through talks, blog posts, and contributions to the open agent ecosystem.
Who You Are:
- You have a BS/MS/PhD in a scientific field, or equivalent experience.
- 10+ years of relevant engineering or applied science experience, including time as a technical lead.
- Proven track record of leading ML or GenAI initiatives in a product-driven environment, from research through production.
- Significant experience with evaluation, experimentation, or measurement of ML systems at scale.
- You bring a strong product mindset and are comfortable driving initiatives across cross-functional teams.
- You thrive in ambiguity and can make sound technical calls when the path isn’t yet defined.
Benefits and Growth:
- New hire stock equity (RSUs) and employee stock purchase plan (ESPP)
- Continuous professional development, product training, and career pathing
- An inclusive company culture, giving programs, and the ability to join our Community Guilds (Datadog employee resource groups)
- Competitive global benefits and global Spring Health benefits for employees and dependents age 6+
- #LI-OnsiteDatadog offers a competitive salary and equity package, and may include variable compensation. Actual compensation is based on factors such as the candidate's skills, qualifications, and experience. In addition, Datadog offers a wide range of best in class, comprehensive and inclusive employee benefits for this role including healthcare, dental, parental planning, and mental health benefits, a 401(k) plan and match, paid time off, fitness reimbursements, and a discounted employee stock purchase plan.
The reasonably estimated yearly salary for this role at Datadog is: $276,000—$345,000 USD
Skills
Ml Systems, Generative AI, Evaluation, Experimentation, Measurement, Applied Research, Technical Leadership, Product Mindset
Similar jobs
AI Research jobsLeads design, implementation, integration, and field validation of tactical autonomy software for unmanned systems and multi-agent missions. Requires extensive autonomy or robotics experience, strong C++ and Python skills, and the ability to obtain a SECRET clearance.
Leads technical direction and develops maritime autonomy capabilities for unmanned surface and underwater vehicles, including motion planning, localization, safe behaviors, and heterogeneous multi-agent collaboration. Requires deep robotics and unmanned-systems experience, strong C++/Python skills, and senior technical leadership.
Research and evaluate frontier AI capabilities for cybersecurity, rapidly prototyping tools, designing rigorous benchmarks, and helping operationalize reliable capabilities into products. Requires deep security expertise, strong technical communication, and at least seven years of relevant experience.
Evaluates model and Generative AI risks across Upstart Bank’s model inventory, conducting risk assessments, monitoring reviews, quantitative analyses, and governance activities. Requires a quantitative master’s degree, 4+ years of relevant experience, and coding skills in Python, R, or similar languages.
Own the architecture, delivery, evaluation, and production operations of AI capabilities embedded in procurement and finance workflows. The role requires 10+ years in applied AI or machine learning, deep LLM and agent expertise, and experience delivering measurable production outcomes.