“It Became My Buddy, But I’m Not Afraid to Disagree”: A Multi-Session Study of UX Evaluators Collaborating with Conversational AI Assistants
Authors
Paper Title
“It Became My Buddy, But I’m Not Afraid to Disagree”: A Multi-Session Study of UX Evaluators Collaborating with Conversational AI Assistants
Publication Info
- Topic area: Human-AI collaboration in usability analysis
- Keywords: UX evaluation, conversational AI, novice-expert paradigm, human-AI collaboration, usability testing, longitudinal study, trust in AI, AI expertise, evaluator behaviors, AI-assisted decision-making
Background and Problem
- Problem / challenge: Fully automated usability analysis often misses complex or subtle issues requiring human contextual understanding, and prior studies of human-AI collaboration have been short-term, neglecting how trust and behaviors evolve over time or how perceived AI expertise impacts collaboration.
- Significance: Understanding how UX evaluators adapt to AI tools and how AI expertise influences usability analysis can improve the effectiveness and reliability of human-AI collaboration in real-world settings.
- Motivation and related work: Previous research has explored AI-assisted usability analysis using visualizations and conversational assistants (CAs), but it has not addressed longitudinal dynamics or the effects of varying AI expertise. This study builds on the novice-expert framework to investigate these gaps.
Solution
- Proposed approach: Development of a usability analysis tool with two custom conversational AI assistants (novice and experienced) to simulate varying levels of UX expertise, evaluated through a multi-session within-subjects study.
- Novelty:
- Longitudinal study design capturing changes in evaluator behaviors and attitudes over five sessions.
- Comparison of novice and experienced CAs to assess the impact of perceived AI expertise on usability analysis.
- Integration of automatic suggestions and reactive responses tailored to UX evaluation tasks.
- Procedure and key techniques:
- Simulated novice and experienced CAs using GPT-4V with tailored prompts reflecting different levels of UX expertise.
- Conducted a within-subjects study with 12 UX evaluators analyzing 15 usability videos across three conditions (no CA, novice CA, experienced CA) over five sessions.
- Collected behavioral data (e.g., video playback strategies, suggestion acceptance rates), analytic performance metrics (e.g., problem coverage, inter-rater reliability), and subjective feedback (e.g., trust, perceived efficiency).
Results
- Concrete findings:
- Experienced CA achieved higher precision (93%), recall (76%), and response accuracy (88%) compared to the novice CA (precision: 80%, recall: 54%, accuracy: 73%).
- Participants identified more usability problems with the experienced CA (8.2 problems/video) than with the novice CA (6.6 problems/video) or no CA (5.5 problems/video).
- Inter-rater reliability (Fleiss’ kappa) was highest with the experienced CA (0.62), followed by the novice CA (0.58) and no CA (0.40).
- Advantage over baselines:
- Both CAs significantly outperformed the no-CA condition in terms of problem identification, coverage, and reliability.
- The experienced CA was rated significantly higher in efficiency, trustworthiness, and completeness of suggestions compared to the novice CA.
- Experiments / evaluation:
- Study design: Within-subjects, five sessions per participant, three conditions (no CA, novice CA, experienced CA).
- Metrics: Precision, recall, response accuracy, problem coverage, Fleiss’ kappa, Likert-scale ratings for trust, efficiency, and confidence.
- Dataset: 15 usability videos covering three products (desktop website, smartphone app, VR headset).
- Limitations and future work:
- Limited sample size (12 participants) and potential gender bias (10 females, 2 males).
- Need for further research on how novice CAs can improve evaluator skills and how personality traits of CAs influence collaboration.
- Future studies could explore multi-CA collaboration and adaptive AI systems that evolve with user expertise.
Summary
This study investigated how UX evaluators adapt to AI tools and how perceived AI expertise affects usability analysis through a multi-session study with novice and experienced conversational AI assistants. Results showed that the experienced CA significantly improved problem identification, trust, and efficiency compared to the novice CA and no-CA conditions. Participants’ behaviors evolved over time, shifting from two-pass to one-pass strategies, with trust and efficiency recovering after an initial novelty effect. The findings highlight the potential of experienced CAs to enhance usability evaluations and the value of novice CAs as training aids. Future work should explore adaptive and multi-CA systems to further optimize human-AI collaboration.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 78%
Interview-Informed Generative Agents for Product Discovery: A Validation Study
CHI '26· Human-LLM Collaboration +3
- 78%
Just-In-Time Objectives: A General Approach for Specialized AI Interactions
CHI '26· Human-LLM Collaboration +3
- 78%
Mapping the Design Space of User Experience for Computer Use Agents
IUI '26· Human-LLM Collaboration +4
- 75%
Does My Chatbot Have an Agenda? Understanding Human and AI Agency in Human-Human-like Chatbot Interaction
CHI '26· Agent Personality & Anthropomorphism +2
- 75%
DuetUI: A Bidirectional Context Loop for Human-Agent Co-Generation of Task-Oriented Interfaces
CHI '26· Human-LLM Collaboration +2
- 75%
Vibe Check: Understanding the Effects of LLM-Based Conversational Agents' Personality and Alignment on User Perceptions in Goal-Oriented Tasks
CHI '26· Agent Personality & Anthropomorphism +2
- 75%
Teaching-Learning Interaction: A New Concept for Interaction Design to Support Reflective User Agency in Intelligent Systems
DIS '21· Human-LLM Collaboration +2
- 67%
DALL: Data Labeling via Data Programming and Active Learning Enhanced by Large Language Models
CHI '26· Human-LLM Collaboration +3
- 67%
Perspectra: Choosing Your Experts Enhances Critical Thinking in Multi-Agent Research Ideation
CHI '26· Human-LLM Collaboration +3
- 67%
"Here, Let Me Help": An Empirical Study of User Interventions in Human–Web Agent Collaboration
CHI '26· Human-LLM Collaboration +3
Based on Jaccard similarity of research subtopics & professions (≥60%)