Human Evaluation Researcher
Designs and operates human-evaluation studies for real-time AI avatars, translating subjective qualities such as naturalness, emotion, trust, and presence into reliable signals for model development and release decisions. Requires 5+ years of qualitative and quantitative human-subjects research experience.
About the job
Responsibilities
- Design and run qualitative and quantitative human-subjects studies of AI avatars, including side-by-side comparisons, controlled rating experiments, interviews, think-alouds, diary studies, and longitudinal panels.
- Convert ambiguous judgments such as naturalness, emotional resonance, trust, and presence into reliable instruments, including rubrics, anchored scales, and annotation guidelines.
- Measure and improve inter-rater agreement while preserving meaningful signal.
- Use ethnographic methods such as observation of live conversations, contextual inquiry, and fieldwork to understand user experiences.
- Build the human-evaluation pipeline, including participant panels, rater training and calibration, tooling, and evaluation cadence tied to model releases.
- Calibrate automated and model-based metrics, including LLM-as-judge systems, against human judgment.
- Work directly with founders and modeling researchers to determine which models ship and what to train next.
Requirements
- 5+ years designing and running human-subjects research in industry or academia, such as UX research, HCI, experimental psychology, behavioral science, or a related field.
- Demonstrated experience designing studies that produce aligned judgments about ambiguous qualities such as tone, emotion, quality, or trust.
- Strong qualitative research skills, including interviews, ethnography, and contextual inquiry.
- Strong quantitative research skills, including survey and psychometric design, experimental design, and statistics for rating and pairwise-comparison data.
- Expertise in agreement and reliability measures, including Cohen’s kappa and Krippendorff’s alpha.
- Ability to run rigorous, practical studies quickly and communicate findings clearly to ML researchers.
Nice-to-haves
- Experience evaluating generative AI, including avatars, digital humans, speech or video generation, conversational agents, or emotion expression and recognition.
- Background in perceptual science or psychophysics.
- MS or PhD in a related field.
- Experience with human evaluation at scale, including crowdsourcing platforms, annotation tooling, and golden datasets.
- Ability to analyze data with Python or R.
Compensation and Benefits
- $160,000–$190,000 base salary plus meaningful equity.
- Health plans, including an HDHP with approximately $2,000 in annual company HSA contributions.
- 15 days of PTO, 10 public holidays, and a full office closure at year-end.
- Weekday lunch, drinks, and snacks.
- Commuter benefits of up to $340 per month.
- 401(k) match.
Skills
Qualitative Research, Quantitative Research, Psychometrics, Experimental Design, Statistics, Cohen'S Kappa, Krippendorff'S Alpha, Ethnography, Contextual Inquiry, Crowdsourcing, Python, R, Llm-As-Judge
Similar jobs
UX Research jobsResearch Manager who partners with customers to design and execute qualitative and quantitative studies, synthesize findings, and translate feedback into actionable business direction. Requires 5+ years in research, insights, or strategy and strong client communication skills.
Leads autonomous qualitative, quantitative, and mixed-methods research to improve the patient journey from healthcare discovery through first-session attendance. Partners with Product, Design, Data/Analytics, and Marketing to turn patient insights into measurable acquisition, onboarding, conversion, and experience improvements.
Leads qualitative and mixed-methods research shaping Reddit’s core consumer experiences, partnering with product, design, engineering, and data science teams. Requires a senior product researcher with 6+ years of consumer UX research experience and strong strategic influence.
Senior user researcher owning mixed-method research for a core product area, from framing and study design through synthesis and product activation. The role requires 6+ years of experience, strong cross-functional influence, and the ability to apply rigorous, AI-supported research practices.
Leads quantitative UX research for high-stakes product and AI investment decisions, developing measurement frameworks, scaled usability studies, and customer segmentation. Requires 8+ years of experience, executive influence, strong statistical expertise, and proficiency in SQL and Python.