Heuristic Evaluation of Conversational Agents

Best Paper
Conversational ChatbotsAgent Personality & Anthropomorphism

Title of the Paper

Heuristic Evaluation of Conversational Agents

Paper Information

  • Subject Area: Usability evaluation in human-computer interaction and conversational agent design
  • Keywords: Conversational agents, usability heuristics, user interface design, voice assistants, chatbots, usability testing, human-computer interaction, data privacy, conversational content, interaction design

Research Background and Problem Statement

  • Identified Problems or Challenges:

    • The current design of conversational agents (e.g., chatbots and voice assistants) lacks consistent usability guidelines.
    • Nielsen's classic usability heuristics have not been effectively validated or adapted for the unique characteristics of conversational agents.
    • Conversational agents differ from traditional human-computer interfaces, involving new design domains such as conversational content, interaction design, privacy, and humanization features.
  • Significance:

    • Ensuring high usability in the design of conversational agents is crucial as they become increasingly popular in commercial and user contexts.
    • Usability heuristic guidelines tailored to different conversational modes (text, voice, or multimodal) can significantly enhance design quality.
  • Research Motivation and Related Work:

    • The authors aim to develop a comprehensive set of usability heuristics for conversational agents, based on an extension of Nielsen's heuristics and expert feedback.
    • The study draws on existing research in human-computer interaction, humanized chatbot experiences, and domain-specific heuristic design.

Proposed Solution

  • Method/Approach:

    • The authors propose a set of 11 usability heuristics specifically designed for conversational agents, validated through four stages:
      1. Heuristic Generation: Synthesizing guidelines from literature and prior studies to establish initial heuristics.
      2. Expert Review: Gathering feedback through expert evaluations and iterating on the initial heuristics.
      3. Heuristic Evaluation Validation: Comparing the effectiveness of the initial heuristics and Nielsen's heuristics in evaluating voice assistants (Amazon Echo Skill) and chatbots.
      4. Revised Heuristic Validation: Refining the heuristics based on evaluation results and conducting further assessments on chatbots.
  • Innovative Contributions:

    • The study proposes and validates specialized heuristic guidelines for conversational systems, extending the traditional heuristic evaluation framework.
    • The guidelines incorporate theories such as Grice's Cooperative Principle to improve the user conversational experience.
    • Explicitly addresses emerging multimodal interaction interface needs, including features like privacy and context maintenance.
  • Implementation Steps and Key Techniques:

    • Initial heuristics were extracted, organized, and categorized from the literature and integrated with Grice's principles to address conversational quality issues.
    • Expert feedback was collected via online surveys, and the heuristics were revised based on rationality ratings.
    • Evaluation stages were systematically designed, mapping heuristics to specific usability issues.

Research Outcomes

  • Specific Results:

    • Developed and validated 11 usability heuristics for conversational agents, covering key design areas such as conversational content, guidance and help, and context maintenance.
    • The revised heuristics outperformed Nielsen's general guidelines in adaptability, identifying more issues related to conversational content and privacy.
  • Advantages Over Existing Solutions:

    • Identified more low-to-medium severity usability issues, particularly in conversational content and visual design domains.
    • Particularly suitable for evaluating interaction designs in multimodal contexts involving voice and text.
  • Experimental or Evaluation Results:

    • The revised heuristics received higher overall relevance ratings from the expert group, identifying significantly more issues in chatbots and voice assistants compared to Nielsen's heuristics.
    • Quantitative analysis further demonstrated the new heuristics' advantage in covering "unique issues."
    • However, the revised heuristics still require improvement in areas such as settings and audio output.
  • Limitations and Future Directions:

    • The study was limited by a small participant group and a narrow range of conversational agent types (only voice assistants and text-based chatbots).
    • Future research is recommended to expand the heuristics' applicability to a broader range of use case scenarios, device types, and domain-specific practices.
    • Further validation is needed to assess the generalizability of the heuristics across diverse user demographics (e.g., different cultural backgrounds or skill levels).

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/47776/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3411764.3445312
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2021
emoji_events
Award
Best Paper
group
Authors
6 authors
sell
Subtopics
Conversational Chatbots, Agent Personality & Anthropomorphism
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
10 related papers