Skip to content
NotionNotion

User Researcher, AI Evaluations

UX Researcher defining and scaling evaluation frameworks for Notion's AI-powered experiences. Focuses on establishing rubrics for model output quality and end-to-end user experience, running longitudinal studies, and operationalizing evaluation processes with product and data science teams.

About the job

What You'll Achieve

Define what “good” looks like (frameworks & rubrics)

  • Establish clear, reusable evaluation criteria that reflect real user expectations—helpfulness, trust, tone, control, and transparency
  • Translate qualitative insight into scoring guidance that can be applied consistently across teams and over time

Run recurring evals (longitudinal & feature-specific)

  • Run recurring longitudinal and feature-specific surveys and studies to measure experience quality over time against defined rubrics
  • Lead qualitative studies, side-by-side comparisons, and human-in-the-loop evaluation efforts
  • Help teams spot regressions, benchmark improvements, and understand when expectations shift

Anchor evaluation in real workflows (context > isolated feedback)

  • Ensure evals reflect jobs-to-be-done, user intent, and the full interaction journey (goal setting, delegation, review, iteration)
  • Help teams understand who is evaluating, what they’re trying to do, and why outputs succeed or fail

Identify failure modes & recovery behavior (guardrails)

  • Uncover breakdowns, regressions, and edge cases across the system—from model behavior to UI and integrations
  • Study how people notice issues, correct them, and continue their work
  • Turn insights into actionable guidance for guardrails, fixes, and prioritization

Operationalize evaluation with partners (process & tooling)

  • Collaborate closely with Product, Design, Engineering, and Data Science to align on target use cases
  • Build scalable evaluation loops (human-in-the-loop review, longitudinal studies, and calibration of automated/LLM-judge approaches against human judgment)

Skills You'll Need to Bring

  • Ability to operationalize insight into measurement: turning “soft” user expectations (trust, tone, usefulness, clarity) into concrete rubrics, scoring guidelines, and observable metrics
  • AI fluency and systems thinking: hands-on with AI products, reasoning about how model behavior, uncertainty, and system constraints shape user experience
  • Experience evaluating AI-enabled products (LLMs, agents, generative UI/workflow automation) and working with Data Science/ML partners on measurement strategy and evaluation tooling
  • Clear communication and impact orientation: aligning diverse partners around shared definitions of quality, creating artifacts that enable teams to act consistently
  • Strong UX research craft (quant + qual): choosing the right methods for the question—interviews, benchmarking, surveys, experiments—and synthesizing into actionable guidance
  • Pragmatism in fast-moving environments: prioritizing ruthlessly, working through ambiguity, balancing scrappy iteration with deep dives

Experience

  • 5+ years doing UX research in industry

Nice to Haves

  • Familiarity with LLM-as-judge methods, prompt design for evaluators, or “golden dataset” creation
  • Experience using AI research tooling for rapid synthesis and communication (e.g., Dovetail, Listen Labs, Maze, Outset, etc.), as well as AI observability tooling like Braintrust
  • Experience using data querying languages (e.g., SQL), scripting languages (e.g., Python), or statistical/mathematical software (e.g., R, SAS, Matlab, etc.)
  • Master’s or PhD in HCI, Psychology, Behavioral Science, Anthropology, Sociology, or a related field

Skills

Ux Research, Qualitative Research, Quantitative Research, Ai Product Evaluation, Llm Evaluation, Rubric Development, Human-In-The-Loop Evaluation, Data Science Collaboration, Python, SQL

Nuance Labs

Nuance Labs

Seattle, WA

Human Evaluation Researcher
$160k+/yrOn-site5+ YOEUX Research

Designs and operates human-evaluation studies for real-time AI avatars, translating subjective qualities such as naturalness, emotion, trust, and presence into reliable signals for model development and release decisions. Requires 5+ years of qualitative and quantitative human-subjects research experience.

Outset

Outset

San Francisco, CA
Research Manager
$150k+/yrHybrid5+ YOEUX Research

Research Manager who partners with customers to design and execute qualitative and quantitative studies, synthesize findings, and translate feedback into actionable business direction. Requires 5+ years in research, insights, or strategy and strong client communication skills.

Gusto

Gusto

Denver, CO
UX Researcher, Quantitative
$172k+/yrHybrid8+ YOEUX Research

Leads quantitative UX research for high-stakes product and AI investment decisions, developing measurement frameworks, scaled usability studies, and customer segmentation. Requires 8+ years of experience, executive influence, strong statistical expertise, and proficiency in SQL and Python.

Reddit

Reddit

United States

Senior UX Researcher, Consumer Product
$168k+/yrRemote6+ YOEUX Research

Leads qualitative and mixed-methods research shaping Reddit’s core consumer experiences, partnering with product, design, engineering, and data science teams. Requires a senior product researcher with 6+ years of consumer UX research experience and strong strategic influence.

Rula

Rula

United States

Senior UX Researcher - Patient
$166k+/yrRemote5+ YOEUX Research

Leads autonomous qualitative, quantitative, and mixed-methods research to improve the patient journey from healthcare discovery through first-session attendance. Partners with Product, Design, Data/Analytics, and Marketing to turn patient insights into measurable acquisition, onboarding, conversion, and experience improvements.