EvAlignUX: Advancing UX Evaluation through LLM-Supported Metrics Exploration
Authors
Research Background and Problem
-
Challenges Identified by the Authors:
User experience (UX) evaluation demonstrates significant shortcomings when dealing with complexity, unpredictability, and the generative capabilities of artificial intelligence (AI). Traditional UX evaluation methods, such as the System Usability Scale (SUS) and the User Experience Questionnaire (UEQ), fail to adequately capture the dynamic, social, and multimodal behaviors inherent in human-computer interaction, particularly in human-AI interaction. -
Why It Matters:
The role of AI systems in user interaction is becoming increasingly significant, yet there remains a substantial gap in assessing whether these systems meet user needs and leveraging their unique characteristics to improve systems. A better understanding of the relationship between evaluation metrics and expected research outcomes can enhance user trust, engagement, and the social impact of systems. -
Research Motivation:
There is currently a lack of tools to assist researchers in selecting appropriate UX evaluation metrics and formulating comprehensive plans. Despite the availability of numerous evaluation methods, researchers still face challenges in scientifically choosing from a plethora of methods and metrics.
Solution
-
Proposed Solution by the Authors:
The authors developed EvAlignUX, an interactive tool supported by large language models (LLM), designed to help UX researchers explore evaluation metrics and their relevance to research outcomes. -
Innovative Features:
- EvAlignUX offers three main functional panels: the Project Ideation Panel, the Metrics Explorer Panel, and the Outcomes and Risks Panel.
- It integrates UX evaluation metrics and related literature into an interactive system, enabling researchers to gain insights from existing studies and make more informed decisions when selecting evaluation metrics.
- The system provides visualized metric graphs to help researchers uncover potential connections between metrics.
-
Implementation Steps and Key Technologies:
- Users input project descriptions, preliminary evaluation plans, and expected outcomes.
- EvAlignUX utilizes GPT-4-based generative technology and a knowledge graph (supported by Neo4j) to recommend relevant evaluation metrics and literature.
- The system offers diverse evaluation views, such as list views and chart views, through a query-based and interactive interface.
- It integrates an AI risk case database to provide users with potential risk warnings and reference suggestions for research outcomes.
Research Outcomes
-
Specific Results:
- EvAlignUX improved users' clarity, completeness, and confidence in UX evaluation planning.
- The system guided participants to think more deeply about the potential impacts and risks of their research, forming a "UX Question Bank" to support future development.
-
Comparison with Existing Solutions and Advantages:
- Compared to traditional methods, EvAlignUX dynamically adjusts users' metric selection, addressing researchers' lack of relevant literature and experience.
- Enhanced features such as graphical interactive interfaces and citation links simplify the metric exploration process, significantly improving efficiency.
- The system helps participants identify overlooked dimensions, such as ethical risks and the generalizability of research outcomes.
-
Experiments and Evaluation:
- Multi-stage user testing involving 19 UX researchers demonstrated that EvAlignUX increased perceived quality of planning, such as a 42% increase in plan refinement and risk consideration.
- Features like the metrics list, charts, and expected outcomes module were widely recognized as useful.
-
Limitations and Future Directions:
- The current knowledge base is limited, covering only selected human-AI interaction research topics and not yet fully encompassing multidisciplinary or industry-specific cases.
- While the system is highly effective for beginners, its role for experienced researchers leans more toward validation rather than inspiration. Future improvements should include more granular adaptive functionalities.
- Further research is needed on the risks of potential dependency on AI systems, such as mitigating dependency through trigger-based question prompts or dynamic feedback mechanisms.
By incorporating LLM technology, EvAlignUX significantly improves the way UX researchers formulate plans, achieving breakthroughs in efficiency, cognitive support, and proactive risk consideration. Its proposed shift from "method-centered" to "thought-centered" evaluation design offers new perspectives for UX assessment and education. Future improvements, such as expanding industrial applications and enhancing self-critical capabilities, will enable its value to be realized across broader contexts.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can traditional UX evaluation methods be improved to better accommodate complexity and multimodal behavior in AI interaction?Category: Visualization Evaluation Methods and Empirical User StudiesSimilar questionsarrow_forward
- Can interactive tools such as EvAlignUX help UX researchers more efficiently select evaluation metrics and develop evaluation plans?Category: Visualization Evaluation Methods and Empirical User StudiesSimilar questionsarrow_forward
- How can associations between UX evaluation metrics and expected research outcomes be visualized and effectively used by researchers?Category: Visualization Evaluation Methods and Empirical User StudiesSimilar questionsarrow_forward
Practical Problems
1- Researchers struggle to scientifically select UX evaluation metrics and cannot comprehensively develop evaluation plans.Category: Visualization Evaluation Methods and Empirical User StudiesSimilar questionsarrow_forward
- 80%
"Why is 'Chicago' deceptive?" Towards Building Model-Driven Tutorials for Humans
CHI '20· Human-LLM Collaboration +2
- 80%
Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies
CHI '25· Human-LLM Collaboration +1
- 75%
COGAM: Measuring and Moderating Cognitive Load in Machine Learning Model Explanations
CHI '20· Explainable AI (XAI)
- 75%
Continual Human-in-the-Loop Optimization
CHI '25· Human-LLM Collaboration
- 67%
The Metacognitive Demands and Opportunities of Generative AI
CHI '24· Generative AI (Text, Image, Music, Video) +2
- 67%
Human Creativity in the Age of LLMs: Randomized Experiments on Divergent and Convergent Thinking
CHI '25· Human-LLM Collaboration +2
- 67%
Simulacrum of stories: Examining Large Language Models as Qualitative Research Participants
CHI '25· Human-LLM Collaboration +2
- 67%
Tell Me What I Missed: Interacting with GPT during Recalling of One-Time Witnessed Events
CHI '26· Human-LLM Collaboration +2
- 67%
A Survey of Collaborative Reinforcement Learning: Interactive Methods and Design Patterns
DIS '21· Human-LLM Collaboration +2
- 67%
PersonaFlow: Designing LLM-Simulated Expert Perspectives for Enhanced Research Ideation
DIS '25· Generative AI (Text, Image, Music, Video) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)