Heuristic Evaluation of Conversational Agents
Best PaperAuthors
Title of the Paper
Heuristic Evaluation of Conversational Agents
Paper Information
- Subject Area: Usability evaluation in human-computer interaction and conversational agent design
- Keywords: Conversational agents, usability heuristics, user interface design, voice assistants, chatbots, usability testing, human-computer interaction, data privacy, conversational content, interaction design
Research Background and Problem Statement
-
Identified Problems or Challenges:
- The current design of conversational agents (e.g., chatbots and voice assistants) lacks consistent usability guidelines.
- Nielsen's classic usability heuristics have not been effectively validated or adapted for the unique characteristics of conversational agents.
- Conversational agents differ from traditional human-computer interfaces, involving new design domains such as conversational content, interaction design, privacy, and humanization features.
-
Significance:
- Ensuring high usability in the design of conversational agents is crucial as they become increasingly popular in commercial and user contexts.
- Usability heuristic guidelines tailored to different conversational modes (text, voice, or multimodal) can significantly enhance design quality.
-
Research Motivation and Related Work:
- The authors aim to develop a comprehensive set of usability heuristics for conversational agents, based on an extension of Nielsen's heuristics and expert feedback.
- The study draws on existing research in human-computer interaction, humanized chatbot experiences, and domain-specific heuristic design.
Proposed Solution
-
Method/Approach:
- The authors propose a set of 11 usability heuristics specifically designed for conversational agents, validated through four stages:
- Heuristic Generation: Synthesizing guidelines from literature and prior studies to establish initial heuristics.
- Expert Review: Gathering feedback through expert evaluations and iterating on the initial heuristics.
- Heuristic Evaluation Validation: Comparing the effectiveness of the initial heuristics and Nielsen's heuristics in evaluating voice assistants (Amazon Echo Skill) and chatbots.
- Revised Heuristic Validation: Refining the heuristics based on evaluation results and conducting further assessments on chatbots.
- The authors propose a set of 11 usability heuristics specifically designed for conversational agents, validated through four stages:
-
Innovative Contributions:
- The study proposes and validates specialized heuristic guidelines for conversational systems, extending the traditional heuristic evaluation framework.
- The guidelines incorporate theories such as Grice's Cooperative Principle to improve the user conversational experience.
- Explicitly addresses emerging multimodal interaction interface needs, including features like privacy and context maintenance.
-
Implementation Steps and Key Techniques:
- Initial heuristics were extracted, organized, and categorized from the literature and integrated with Grice's principles to address conversational quality issues.
- Expert feedback was collected via online surveys, and the heuristics were revised based on rationality ratings.
- Evaluation stages were systematically designed, mapping heuristics to specific usability issues.
Research Outcomes
-
Specific Results:
- Developed and validated 11 usability heuristics for conversational agents, covering key design areas such as conversational content, guidance and help, and context maintenance.
- The revised heuristics outperformed Nielsen's general guidelines in adaptability, identifying more issues related to conversational content and privacy.
-
Advantages Over Existing Solutions:
- Identified more low-to-medium severity usability issues, particularly in conversational content and visual design domains.
- Particularly suitable for evaluating interaction designs in multimodal contexts involving voice and text.
-
Experimental or Evaluation Results:
- The revised heuristics received higher overall relevance ratings from the expert group, identifying significantly more issues in chatbots and voice assistants compared to Nielsen's heuristics.
- Quantitative analysis further demonstrated the new heuristics' advantage in covering "unique issues."
- However, the revised heuristics still require improvement in areas such as settings and audio output.
-
Limitations and Future Directions:
- The study was limited by a small participant group and a narrow range of conversational agent types (only voice assistants and text-based chatbots).
- Future research is recommended to expand the heuristics' applicability to a broader range of use case scenarios, device types, and domain-specific practices.
- Further validation is needed to assess the generalizability of the heuristics across diverse user demographics (e.g., different cultural backgrounds or skill levels).
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can a set of applicable usability heuristic guidelines be designed for conversational agents (e.g., chatbots and voice assistants)?Category: VUI Design Methods, Guidelines, and Heuristic EvaluationSimilar questionsarrow_forward
- What advantages and disadvantages do these specifically designed heuristic guidelines have compared to Nielsen's classic heuristics?Category: VUI Design Methods, Guidelines, and Heuristic EvaluationSimilar questionsarrow_forward
- How can the validity of these heuristic guidelines in multimodal interaction scenarios be verified?Category: VUI Design Methods, Guidelines, and Heuristic EvaluationSimilar questionsarrow_forward
Practical Problems
1- Users may encounter insufficiently smooth interaction or privacy issues when using chatbots or voice assistants.Category: VUI Design Methods, Guidelines, and Heuristic EvaluationSimilar questionsarrow_forward
- 100%
Touch Your Heart: A Tone-aware Chatbot for Customer Care on Social Media
CHI '18· Conversational Chatbots +1
- 100%
Single or Multiple Conversational Agents? An Interactional Coherence Comparison
CHI '18· Conversational Chatbots +1
- 100%
What Makes a Good Conversation? Challenges in Designing Truly Conversational Agents
CHI '19· Conversational Chatbots +1
- 100%
If I Hear You Correctly: Building and Evaluating Interview Chatbots with Active Listening Skills
CHI '20· Conversational Chatbots +1
- 100%
Bot in the Bunch: Facilitating Group Chat Discussion by Improving Efficiency and Participation with a Chatbot
CHI '20· Conversational Chatbots +1
- 100%
Effects of Persuasive Dialogues: Testing Bot Identities and Inquiry Strategies
CHI '20· Conversational Chatbots +1
- 100%
"I Hear You, I Feel You": Encouraging Deep Self-disclosure through a Chatbot
CHI '20· Conversational Chatbots +1
- 100%
Exploring Semi-Supervised Learning for Predicting Listener Backchannels
CHI '21· Conversational Chatbots +1
- 100%
Designing Conversational Agents: A Self-Determination Theory Approach
CHI '21· Conversational Chatbots +1
- 100%
Collaborating with a Text-Based Chatbot: An Exploration of Real-World Collaboration Strategies Enacted during Human-Chatbot Interactions
CHI '23· Conversational Chatbots +1
Based on Jaccard similarity of research subtopics & professions (≥60%)