UICrit: Enhancing Automated Design Evaluation with a UI Critique Dataset
Authors
Automated UI evaluation can be beneficial for the design process; for example, to compare different UI designs, or conduct automated heuristic evaluation. LLM-based UI evaluation, in particular, holds the promise of generalizability to a wide variety of UI types and evaluation tasks. However, current LLM-based techniques do not yet match the performance of human evaluators. We hypothesize that automatic evaluation can be improved by collecting a targeted UI feedback dataset and then using this dataset to enhance the performance of general-purpose LLMs. We present a targeted dataset of 3,059 design critiques and quality ratings for 983 mobile UIs, collected from seven designers, each with at least a year of professional design experience. We carried out an in-depth analysis to characterize the dataset's features. We then applied this dataset to achieve a 55\% performance gain in LLM-generated UI feedback via various few-shot and visual prompting techniques. We also discuss future applications of this dataset, including training a reward model for generative UI techniques, and fine-tuning a tool-agnostic multi-modal LLM that automates UI evaluation.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 80%
Designerly Understanding: Information Needs for Model Transparency to Support Design Ideation for AI-Powered User Experience
CHI '23· Human-LLM Collaboration +2
- 80%
ONYX: Assisting Users in Teaching Natural Language Interfaces Through Multi-Modal Interactive Task Learning
CHI '23· Voice User Interface (VUI) Design +2
- 80%
ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis Testing
CHI '24· Human-LLM Collaboration +2
- 80%
Conversation Progress Guide : UI System for Enhancing Self-Efficacy in Conversational AI
CHI '25· Conversational Chatbots +2
- 80%
Characterizing Unintended Consequences of GUI Agents For Web Browsing
CHI '26· Human-LLM Collaboration +2
- 80%
UIClip: A Data-driven Model for Assessing User Interface Design
UIST '24· 360° Video & Panoramic Content +2
- 75%
Cells, Generators, and Lenses: Design Framework for Object-Oriented Interaction with Large Language Models
UIST '23· Human-LLM Collaboration
- 67%
Planning for Natural Language Failures with the AI Playbook
CHI '21· Human-LLM Collaboration +2
- 67%
Adapting User Interfaces with Model-based Reinforcement Learning
CHI '21· Human-LLM Collaboration +2
- 67%
Designing Responsible AI: Adaptations of UX Practice to Meet Responsible AI Challenges
CHI '23· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)