Studying the Effects of Cognitive Biases in Evaluation of Conversational Agents

Honorable Mention
Conversational ChatbotsUser Research Methods (Interviews, Surveys, Observation)HCI ResearchersAmazon Mechanical Turk Workers

Humans quite frequently interact with conversational agents. The rapid advancement in generative language modeling through neural networks has helped advance the creation of intelligent conversational agents. Researchers typically evaluate the output of their models through crowdsourced judgments, but there are no established best practices for conducting such studies. Moreover, it is unclear if cognitive biases in decision-making are affecting crowdsourced workers' judgments when they undertake these tasks. To investigate, we conducted a between-subjects study with 77 crowdsourced workers to understand the role of cognitive biases, specifically anchoring bias, when humans are asked to evaluate the output of conversational agents. Our results provide insight into how best to evaluate conversational agents. We find increased consistency in ratings across two experimental conditions may be a result of anchoring bias. We also determine that external factors such as time and prior experience in similar tasks have effects on inter-rater consistency.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/32314/2020

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3313831.3376318
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2020
emoji_events
Award
Honorable Mention
group
Authors
3 authors
sell
Subtopics
Conversational Chatbots, User Research Methods (Interviews, Surveys, Observation)
work
Professions
HCI Researchers, Amazon Mechanical Turk Workers
article
Content Status
Abstract only
hub
Related Papers
4 related papers