Enhancing UX Evaluation Through Collaboration with Conversational AI Assistants: Effects of Proactive Dialogue and Timing

Human-LLM CollaborationPrototyping & User TestingComputational Methods in HCIUI/UX Designers

Title of the Paper

Collaboration with Conversational AI Assistants to Enhance User Experience (UX) Evaluation: Proactive Dialogues and Temporal Impacts

Bibliographic Information

  • Field of Study: User Experience (UX) Evaluation and Human-AI Collaboration
  • Keywords: User Experience, Usability Testing, Human-AI Collaboration, Conversational Assistants, Proactive Dialogue, Time Optimization

Research Background and Problem Statement

  • Problems and Challenges:

    • Traditional UX testing video analysis is complex and time-consuming, often leading to information omissions or misjudgments of usability issues.
    • While AI-assisted analysis tools have been explored, most rely on user-driven interaction models, where the system only responds when prompted by the user. The potential of proactive system dialogues remains unclear.
  • Significance:

    • Leveraging AI technology to analyze UX testing videos is a potentially efficient approach that could alleviate time and resource constraints in industrial practices.
    • Investigating the proactive dialogue capabilities of conversational AI assistants and identifying the optimal timing for suggestions can guide the design of future tools, enhancing AI-assisted decision-making and collaboration efficiency.
  • Research Motivation:

    • Responding to Jakob Nielsen's call to integrate AI into UX research, this study explores how proactive conversational assistants can support UX evaluation through automated suggestions.
    • The focus is on the impact of suggestion timing (before, during, or after an issue occurs).

Solution

  • Methods and Innovations:

    • Propose and utilize a proactive conversational assistant (CA) that provides automated usability issue suggestions at three different timings: before, concurrently, or after the issue occurs.
    • Employ the ChatGPT model to generate usability issue suggestions, combined with a "Wizard of Oz" experimental design to simulate realistic CA interactions under controlled conditions.
    • Implement a fixed-timing suggestion display strategy to avoid unexpected interruptions, paired with clear suggestion content (e.g., issue descriptions and timestamps).
  • Implementation Steps and Techniques:

    1. Use ChatGPT to generate usability issue suggestions from testing videos, with preliminary quality evaluation by UX experts.
    2. Design a UX analysis tool featuring a video player and chat interface, enabling functionalities such as recording issues and responding to suggestions.
    3. Conduct a controlled experiment with 24 UX evaluators to assess the impact of proactive suggestions at different timings on user behavior and analysis outcomes.
    4. Analyze experimental data using mixed methods (quantitative and qualitative), including precision and recall evaluation of ChatGPT's suggestions and participant feedback surveys.

Research Outcomes

  • Specific Findings:

    • The timing of automated suggestions had no significant impact on analysis performance (e.g., the number of identified issues).
    • Most participants preferred suggestions provided after potential issues occurred ("after" condition), which significantly improved trust and efficiency ratings.
    • 77.6% of ChatGPT's suggestions were accepted by participants, but the model missed 58.8% of the total issues identified by participants.
  • Comparison with Existing Solutions and Advantages:

    • Proactive dialogues outperform user-driven interactions in fostering user trust and providing timely efficiency improvements.
    • ChatGPT's suggestions served as a validation tool for user analyses but did not replace the professional judgment of UX evaluators.
  • Experimental and Evaluation Results:

    • In the "after" condition, participants' acceptance of suggestions was significantly higher than in other conditions, reaching an 88.1% acceptance rate.
    • Typical responses to ChatGPT's suggestions included agreement, correction, seeking clarification, disagreement, or ignoring.
  • Limitations and Future Directions:

    • Limitations:
      1. ChatGPT's analysis was based solely on text transcripts, lacking the ability to capture multimodal cues (e.g., facial expressions, user behaviors) from videos.
      2. The experimental setup may not fully encompass the complexity of all UX testing scenarios.
    • Future Directions:
      1. Integrate multimodal video analysis capabilities into conversational assistants (e.g., combining video imagery and audio data) to capture user behaviors and emotions more comprehensively.
      2. Explore more personalized suggestion timing, such as allowing users to adjust the delay of suggestions.
      3. Conduct longitudinal studies to observe long-term behavioral changes in UX evaluators.
      4. Investigate the transparency of AI capabilities and appropriate human-AI collaboration models to reduce misinformation and imbalance in reliance on professional expertise.

This study provides valuable insights into human-AI collaboration and temporal optimization design, offering inspiration for the development of smarter and more efficient UX evaluation support tools in the future.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147731/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642168
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Human-LLM Collaboration, Prototyping & User Testing, Computational Methods in HCI
work
Professions
UI/UX Designers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers