UICrit: Enhancing Automated Design Evaluation with a UI Critique Dataset

Human-LLM CollaborationExplainable AI (XAI)UI/UX DesignersAI/ML Researchers & Engineers

Automated UI evaluation can be beneficial for the design process; for example, to compare different UI designs, or conduct automated heuristic evaluation. LLM-based UI evaluation, in particular, holds the promise of generalizability to a wide variety of UI types and evaluation tasks. However, current LLM-based techniques do not yet match the performance of human evaluators. We hypothesize that automatic evaluation can be improved by collecting a targeted UI feedback dataset and then using this dataset to enhance the performance of general-purpose LLMs. We present a targeted dataset of 3,059 design critiques and quality ratings for 983 mobile UIs, collected from seven designers, each with at least a year of professional design experience. We carried out an in-depth analysis to characterize the dataset's features. We then applied this dataset to achieve a 55\% performance gain in LLM-generated UI feedback via various few-shot and visual prompting techniques. We also discuss future applications of this dataset, including training a reward model for generative UI techniques, and fine-tuning a tool-agnostic multi-modal LLM that automates UI evaluation.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/uist/170823/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3654777.3676381
At a Glance

Paper Snapshot

fact_check
dataset
Source
UIST
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Human-LLM Collaboration, Explainable AI (XAI)
work
Professions
UI/UX Designers, AI/ML Researchers & Engineers
article
Content Status
Abstract only
hub
Related Papers
10 related papers