“It Became My Buddy, But I’m Not Afraid to Disagree”: A Multi-Session Study of UX Evaluators Collaborating with Conversational AI Assistants

Human-LLM CollaborationAI-Assisted Decision-Making & AutomationAgent Personality & AnthropomorphismUser Research Methods (Interviews, Surveys, Observation)Prototyping & User TestingUI/UX DesignersHCI ResearchersAI/ML Researchers & Engineers

Paper Title

“It Became My Buddy, But I’m Not Afraid to Disagree”: A Multi-Session Study of UX Evaluators Collaborating with Conversational AI Assistants

Publication Info

  • Topic area: Human-AI collaboration in usability analysis
  • Keywords: UX evaluation, conversational AI, novice-expert paradigm, human-AI collaboration, usability testing, longitudinal study, trust in AI, AI expertise, evaluator behaviors, AI-assisted decision-making

Background and Problem

  • Problem / challenge: Fully automated usability analysis often misses complex or subtle issues requiring human contextual understanding, and prior studies of human-AI collaboration have been short-term, neglecting how trust and behaviors evolve over time or how perceived AI expertise impacts collaboration.
  • Significance: Understanding how UX evaluators adapt to AI tools and how AI expertise influences usability analysis can improve the effectiveness and reliability of human-AI collaboration in real-world settings.
  • Motivation and related work: Previous research has explored AI-assisted usability analysis using visualizations and conversational assistants (CAs), but it has not addressed longitudinal dynamics or the effects of varying AI expertise. This study builds on the novice-expert framework to investigate these gaps.

Solution

  • Proposed approach: Development of a usability analysis tool with two custom conversational AI assistants (novice and experienced) to simulate varying levels of UX expertise, evaluated through a multi-session within-subjects study.
  • Novelty:
    1. Longitudinal study design capturing changes in evaluator behaviors and attitudes over five sessions.
    2. Comparison of novice and experienced CAs to assess the impact of perceived AI expertise on usability analysis.
    3. Integration of automatic suggestions and reactive responses tailored to UX evaluation tasks.
  • Procedure and key techniques:
    • Simulated novice and experienced CAs using GPT-4V with tailored prompts reflecting different levels of UX expertise.
    • Conducted a within-subjects study with 12 UX evaluators analyzing 15 usability videos across three conditions (no CA, novice CA, experienced CA) over five sessions.
    • Collected behavioral data (e.g., video playback strategies, suggestion acceptance rates), analytic performance metrics (e.g., problem coverage, inter-rater reliability), and subjective feedback (e.g., trust, perceived efficiency).

Results

  • Concrete findings:
    • Experienced CA achieved higher precision (93%), recall (76%), and response accuracy (88%) compared to the novice CA (precision: 80%, recall: 54%, accuracy: 73%).
    • Participants identified more usability problems with the experienced CA (8.2 problems/video) than with the novice CA (6.6 problems/video) or no CA (5.5 problems/video).
    • Inter-rater reliability (Fleiss’ kappa) was highest with the experienced CA (0.62), followed by the novice CA (0.58) and no CA (0.40).
  • Advantage over baselines:
    • Both CAs significantly outperformed the no-CA condition in terms of problem identification, coverage, and reliability.
    • The experienced CA was rated significantly higher in efficiency, trustworthiness, and completeness of suggestions compared to the novice CA.
  • Experiments / evaluation:
    • Study design: Within-subjects, five sessions per participant, three conditions (no CA, novice CA, experienced CA).
    • Metrics: Precision, recall, response accuracy, problem coverage, Fleiss’ kappa, Likert-scale ratings for trust, efficiency, and confidence.
    • Dataset: 15 usability videos covering three products (desktop website, smartphone app, VR headset).
  • Limitations and future work:
    • Limited sample size (12 participants) and potential gender bias (10 females, 2 males).
    • Need for further research on how novice CAs can improve evaluator skills and how personality traits of CAs influence collaboration.
    • Future studies could explore multi-CA collaboration and adaptive AI systems that evolve with user expertise.

Summary

This study investigated how UX evaluators adapt to AI tools and how perceived AI expertise affects usability analysis through a multi-session study with novice and experienced conversational AI assistants. Results showed that the experienced CA significantly improved problem identification, trust, and efficiency compared to the novice CA and no-CA conditions. Participants’ behaviors evolved over time, shifting from two-pass to one-pass strategies, with trust and efficiency recovering after an initial novelty effect. The findings highlight the potential of experienced CAs to enhance usability evaluations and the value of novice CAs as training aids. Future work should explore adaptive and multi-CA systems to further optimize human-AI collaboration.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222798/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790536
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Decision-Making & Automation, Agent Personality & Anthropomorphism, User Research Methods (Interviews, Surveys, Observation)
work
Professions
UI/UX Designers, HCI Researchers, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers