Enhancing UX Evaluation Through Collaboration with Conversational AI Assistants: Effects of Proactive Dialogue and Timing
Authors
Title of the Paper
Collaboration with Conversational AI Assistants to Enhance User Experience (UX) Evaluation: Proactive Dialogues and Temporal Impacts
Bibliographic Information
- Field of Study: User Experience (UX) Evaluation and Human-AI Collaboration
- Keywords: User Experience, Usability Testing, Human-AI Collaboration, Conversational Assistants, Proactive Dialogue, Time Optimization
Research Background and Problem Statement
-
Problems and Challenges:
- Traditional UX testing video analysis is complex and time-consuming, often leading to information omissions or misjudgments of usability issues.
- While AI-assisted analysis tools have been explored, most rely on user-driven interaction models, where the system only responds when prompted by the user. The potential of proactive system dialogues remains unclear.
-
Significance:
- Leveraging AI technology to analyze UX testing videos is a potentially efficient approach that could alleviate time and resource constraints in industrial practices.
- Investigating the proactive dialogue capabilities of conversational AI assistants and identifying the optimal timing for suggestions can guide the design of future tools, enhancing AI-assisted decision-making and collaboration efficiency.
-
Research Motivation:
- Responding to Jakob Nielsen's call to integrate AI into UX research, this study explores how proactive conversational assistants can support UX evaluation through automated suggestions.
- The focus is on the impact of suggestion timing (before, during, or after an issue occurs).
Solution
-
Methods and Innovations:
- Propose and utilize a proactive conversational assistant (CA) that provides automated usability issue suggestions at three different timings: before, concurrently, or after the issue occurs.
- Employ the ChatGPT model to generate usability issue suggestions, combined with a "Wizard of Oz" experimental design to simulate realistic CA interactions under controlled conditions.
- Implement a fixed-timing suggestion display strategy to avoid unexpected interruptions, paired with clear suggestion content (e.g., issue descriptions and timestamps).
-
Implementation Steps and Techniques:
- Use ChatGPT to generate usability issue suggestions from testing videos, with preliminary quality evaluation by UX experts.
- Design a UX analysis tool featuring a video player and chat interface, enabling functionalities such as recording issues and responding to suggestions.
- Conduct a controlled experiment with 24 UX evaluators to assess the impact of proactive suggestions at different timings on user behavior and analysis outcomes.
- Analyze experimental data using mixed methods (quantitative and qualitative), including precision and recall evaluation of ChatGPT's suggestions and participant feedback surveys.
Research Outcomes
-
Specific Findings:
- The timing of automated suggestions had no significant impact on analysis performance (e.g., the number of identified issues).
- Most participants preferred suggestions provided after potential issues occurred ("after" condition), which significantly improved trust and efficiency ratings.
- 77.6% of ChatGPT's suggestions were accepted by participants, but the model missed 58.8% of the total issues identified by participants.
-
Comparison with Existing Solutions and Advantages:
- Proactive dialogues outperform user-driven interactions in fostering user trust and providing timely efficiency improvements.
- ChatGPT's suggestions served as a validation tool for user analyses but did not replace the professional judgment of UX evaluators.
-
Experimental and Evaluation Results:
- In the "after" condition, participants' acceptance of suggestions was significantly higher than in other conditions, reaching an 88.1% acceptance rate.
- Typical responses to ChatGPT's suggestions included agreement, correction, seeking clarification, disagreement, or ignoring.
-
Limitations and Future Directions:
- Limitations:
- ChatGPT's analysis was based solely on text transcripts, lacking the ability to capture multimodal cues (e.g., facial expressions, user behaviors) from videos.
- The experimental setup may not fully encompass the complexity of all UX testing scenarios.
- Future Directions:
- Integrate multimodal video analysis capabilities into conversational assistants (e.g., combining video imagery and audio data) to capture user behaviors and emotions more comprehensively.
- Explore more personalized suggestion timing, such as allowing users to adjust the delay of suggestions.
- Conduct longitudinal studies to observe long-term behavioral changes in UX evaluators.
- Investigate the transparency of AI capabilities and appropriate human-AI collaboration models to reduce misinformation and imbalance in reliance on professional expertise.
- Limitations:
This study provides valuable insights into human-AI collaboration and temporal optimization design, offering inspiration for the development of smarter and more efficient UX evaluation support tools in the future.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- Can proactive dialogue features in conversational AI assistants improve the efficiency and trust of UX evaluation?Category: Trust in Conversational AI and ChatbotsSimilar questionsarrow_forward
- In UX evaluation, how do automatically provided suggestions at different time points (before, during, and after issues) affect analysis outcomes?Category: Trust in Conversational AI and ChatbotsSimilar questionsarrow_forward
- To what extent are suggestions generated by conversational AI assistants accepted by users and supportive of professional judgment?Category: Trust in Conversational AI and ChatbotsSimilar questionsarrow_forward
Practical Problems
1- Traditional UX testing video analysis is tedious and prone to missing issues.Category: Trust in Conversational AI and ChatbotsSimilar questionsarrow_forward
- 75%
Steering Performance with Error-accepting Delays
CHI '19· Prototyping & User Testing +1
- 75%
SimUser: Generating Usability Feedback by Simulating Various Users Interacting with Mobile Applications
CHI '24· Human-LLM Collaboration +1
- 67%
The Way We Notice, That’s What Really Matters: Instantiating UI Components with Distinguishing Variations
CHI '26· Human-LLM Collaboration +2
- 67%
Athena: Intermediate Representations for Iterative Scaffolded App Generation with an LLM
IUI '26· Human-LLM Collaboration +2
- 60%
Modeling Mobile Interface Tappability Using Crowdsourcing and Deep Learning
CHI '19· Prototyping & User Testing +1
- 60%
Modeling Fully and Partially Constrained Lasso Movements in a Grid of Icons
CHI '19· Prototyping & User Testing +1
- 60%
ORCSolver: An Efficient Solver for Adaptive GUI Layout with OR-Constraints
CHI '20· Prototyping & User Testing +1
- 60%
Investigating the Homogenization of Web Design: A Mixed-Methods Approach
CHI '21· Prototyping & User Testing +1
- 60%
Varv: Reprogrammable Interactive Software as a Declarative Data Structure
CHI '22· Prototyping & User Testing +1
- 60%
Learning to Denoise Raw Mobile UI Layouts for Improving Datasets at Scale
CHI '22· Prototyping & User Testing +1
Based on Jaccard similarity of research subtopics & professions (≥60%)