EvAlignUX: Advancing UX Evaluation through LLM-Supported Metrics Exploration

Human-LLM CollaborationExplainable AI (XAI)HCI ResearchersCognitive Scientists

Research Background and Problem

  • Challenges Identified by the Authors:
    User experience (UX) evaluation demonstrates significant shortcomings when dealing with complexity, unpredictability, and the generative capabilities of artificial intelligence (AI). Traditional UX evaluation methods, such as the System Usability Scale (SUS) and the User Experience Questionnaire (UEQ), fail to adequately capture the dynamic, social, and multimodal behaviors inherent in human-computer interaction, particularly in human-AI interaction.

  • Why It Matters:
    The role of AI systems in user interaction is becoming increasingly significant, yet there remains a substantial gap in assessing whether these systems meet user needs and leveraging their unique characteristics to improve systems. A better understanding of the relationship between evaluation metrics and expected research outcomes can enhance user trust, engagement, and the social impact of systems.

  • Research Motivation:
    There is currently a lack of tools to assist researchers in selecting appropriate UX evaluation metrics and formulating comprehensive plans. Despite the availability of numerous evaluation methods, researchers still face challenges in scientifically choosing from a plethora of methods and metrics.

Solution

  • Proposed Solution by the Authors:
    The authors developed EvAlignUX, an interactive tool supported by large language models (LLM), designed to help UX researchers explore evaluation metrics and their relevance to research outcomes.

  • Innovative Features:

    1. EvAlignUX offers three main functional panels: the Project Ideation Panel, the Metrics Explorer Panel, and the Outcomes and Risks Panel.
    2. It integrates UX evaluation metrics and related literature into an interactive system, enabling researchers to gain insights from existing studies and make more informed decisions when selecting evaluation metrics.
    3. The system provides visualized metric graphs to help researchers uncover potential connections between metrics.
  • Implementation Steps and Key Technologies:

    1. Users input project descriptions, preliminary evaluation plans, and expected outcomes.
    2. EvAlignUX utilizes GPT-4-based generative technology and a knowledge graph (supported by Neo4j) to recommend relevant evaluation metrics and literature.
    3. The system offers diverse evaluation views, such as list views and chart views, through a query-based and interactive interface.
    4. It integrates an AI risk case database to provide users with potential risk warnings and reference suggestions for research outcomes.

Research Outcomes

  • Specific Results:

    1. EvAlignUX improved users' clarity, completeness, and confidence in UX evaluation planning.
    2. The system guided participants to think more deeply about the potential impacts and risks of their research, forming a "UX Question Bank" to support future development.
  • Comparison with Existing Solutions and Advantages:

    1. Compared to traditional methods, EvAlignUX dynamically adjusts users' metric selection, addressing researchers' lack of relevant literature and experience.
    2. Enhanced features such as graphical interactive interfaces and citation links simplify the metric exploration process, significantly improving efficiency.
    3. The system helps participants identify overlooked dimensions, such as ethical risks and the generalizability of research outcomes.
  • Experiments and Evaluation:

    1. Multi-stage user testing involving 19 UX researchers demonstrated that EvAlignUX increased perceived quality of planning, such as a 42% increase in plan refinement and risk consideration.
    2. Features like the metrics list, charts, and expected outcomes module were widely recognized as useful.
  • Limitations and Future Directions:

    1. The current knowledge base is limited, covering only selected human-AI interaction research topics and not yet fully encompassing multidisciplinary or industry-specific cases.
    2. While the system is highly effective for beginners, its role for experienced researchers leans more toward validation rather than inspiration. Future improvements should include more granular adaptive functionalities.
    3. Further research is needed on the risks of potential dependency on AI systems, such as mitigating dependency through trigger-based question prompts or dynamic feedback mechanisms.

By incorporating LLM technology, EvAlignUX significantly improves the way UX researchers formulate plans, achieving breakthroughs in efficiency, cognitive support, and proactive risk consideration. Its proposed shift from "method-centered" to "thought-centered" evaluation design offers new perspectives for UX assessment and education. Future improvements, such as expanding industrial applications and enhancing self-critical capabilities, will enable its value to be realized across broader contexts.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/189073/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3714045
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Human-LLM Collaboration, Explainable AI (XAI)
work
Professions
HCI Researchers, Cognitive Scientists
article
Content Status
Full text indexed
hub
Related Papers
10 related papers