User Researcher, AI Evaluations
UX Researcher defining and scaling evaluation frameworks for Notion's AI-powered experiences. Focuses on establishing rubrics for model output quality and end-to-end user experience, running longitudinal studies, and operationalizing evaluation processes with product and data science teams.
About the job
What You'll Achieve
Define what “good” looks like (frameworks & rubrics)
- Establish clear, reusable evaluation criteria that reflect real user expectations—helpfulness, trust, tone, control, and transparency
- Translate qualitative insight into scoring guidance that can be applied consistently across teams and over time
Run recurring evals (longitudinal & feature-specific)
- Run recurring longitudinal and feature-specific surveys and studies to measure experience quality over time against defined rubrics
- Lead qualitative studies, side-by-side comparisons, and human-in-the-loop evaluation efforts
- Help teams spot regressions, benchmark improvements, and understand when expectations shift
Anchor evaluation in real workflows (context > isolated feedback)
- Ensure evals reflect jobs-to-be-done, user intent, and the full interaction journey (goal setting, delegation, review, iteration)
- Help teams understand who is evaluating, what they’re trying to do, and why outputs succeed or fail
Identify failure modes & recovery behavior (guardrails)
- Uncover breakdowns, regressions, and edge cases across the system—from model behavior to UI and integrations
- Study how people notice issues, correct them, and continue their work
- Turn insights into actionable guidance for guardrails, fixes, and prioritization
Operationalize evaluation with partners (process & tooling)
- Collaborate closely with Product, Design, Engineering, and Data Science to align on target use cases
- Build scalable evaluation loops (human-in-the-loop review, longitudinal studies, and calibration of automated/LLM-judge approaches against human judgment)
Skills You'll Need to Bring
- Ability to operationalize insight into measurement: turning “soft” user expectations (trust, tone, usefulness, clarity) into concrete rubrics, scoring guidelines, and observable metrics
- AI fluency and systems thinking: hands-on with AI products, reasoning about how model behavior, uncertainty, and system constraints shape user experience
- Experience evaluating AI-enabled products (LLMs, agents, generative UI/workflow automation) and working with Data Science/ML partners on measurement strategy and evaluation tooling
- Clear communication and impact orientation: aligning diverse partners around shared definitions of quality, creating artifacts that enable teams to act consistently
- Strong UX research craft (quant + qual): choosing the right methods for the question—interviews, benchmarking, surveys, experiments—and synthesizing into actionable guidance
- Pragmatism in fast-moving environments: prioritizing ruthlessly, working through ambiguity, balancing scrappy iteration with deep dives
Experience
- 5+ years doing UX research in industry
Nice to Haves
- Familiarity with LLM-as-judge methods, prompt design for evaluators, or “golden dataset” creation
- Experience using AI research tooling for rapid synthesis and communication (e.g., Dovetail, Listen Labs, Maze, Outset, etc.), as well as AI observability tooling like Braintrust
- Experience using data querying languages (e.g., SQL), scripting languages (e.g., Python), or statistical/mathematical software (e.g., R, SAS, Matlab, etc.)
- Master’s or PhD in HCI, Psychology, Behavioral Science, Anthropology, Sociology, or a related field
Skills
Ux Research, Qualitative Research, Quantitative Research, Ai Product Evaluation, Llm Evaluation, Rubric Development, Human-In-The-Loop Evaluation, Data Science Collaboration, Python, SQL
Similar jobs
UX Research jobsDesigns and operates human-evaluation studies for real-time AI avatars, translating subjective qualities such as naturalness, emotion, trust, and presence into reliable signals for model development and release decisions. Requires 5+ years of qualitative and quantitative human-subjects research experience.
Research Manager who partners with customers to design and execute qualitative and quantitative studies, synthesize findings, and translate feedback into actionable business direction. Requires 5+ years in research, insights, or strategy and strong client communication skills.
Leads quantitative UX research for high-stakes product and AI investment decisions, developing measurement frameworks, scaled usability studies, and customer segmentation. Requires 8+ years of experience, executive influence, strong statistical expertise, and proficiency in SQL and Python.
Leads qualitative and mixed-methods research shaping Reddit’s core consumer experiences, partnering with product, design, engineering, and data science teams. Requires a senior product researcher with 6+ years of consumer UX research experience and strong strategic influence.
Leads autonomous qualitative, quantitative, and mixed-methods research to improve the patient journey from healthcare discovery through first-session attendance. Partners with Product, Design, Data/Analytics, and Marketing to turn patient insights into measurable acquisition, onboarding, conversion, and experience improvements.