Generating Automatic Feedback on UI Mockups with Large Language Models
Authors
Title of the Paper
Generating Automatic Feedback on UI Mockups with Large Language Models
Paper Information
- Research Area: Human-Computer Interaction, Interface Design, Application of AI in Design
- Keywords: Large Language Models, Human-Computer Interface Design Tools, GPT-4, Heuristic Evaluation, Design Feedback
Research Background and Problem
- Problem or Challenge: In the UI design process, human feedback is a crucial factor for improving designs. However, experienced evaluators are often difficult to access, and manual feedback is costly. Additionally, existing automated feedback tools are limited in evaluation dimensions and struggle to provide interpretable analysis results.
- Significance: UI design directly impacts the quality of user interaction with technology and information. Fast and effective feedback can significantly enhance design efficiency while reducing the subjectivity of design quality.
- Research Motivation: Heuristic evaluation is a widely used design evaluation method but still requires experienced human evaluators to conduct manual checks. The potential of AI-assisted evaluation remains underexplored, particularly for supporting large-scale, automated assessments.
- Related Work: Existing research leverages computer vision and machine learning techniques to provide design feedback, but retraining models for different tasks has limitations. Emerging applications of generative AI demonstrate new possibilities for design support, but their application to heuristic evaluation of UI remains insufficiently explored.
Solution
-
Method or Solution:
- Developed a GPT-4-based large language model tool for automated heuristic evaluation of UI prototypes.
- Implemented the tool as a Figma plugin to evaluate static UI mockups and generate constructive textual suggestions.
- Introduced a JSON-based representation of UI to enable the model to process complex UI hierarchies and semantic information.
- Supported user-defined evaluation guidelines and dynamically adjusted model outputs based on designer feedback.
-
Innovations:
- Cross-guideline support for diverse design feedback.
- Emphasis on presenting feedback in context to enhance designers' understanding of its implications.
- Introduced a reflection mechanism where the model self-corrects based on contextual feedback when designers mark suggestions as incorrect.
-
Implementation Steps and Key Technologies:
- Constructed a JSON representation of UI screens, including semantic features (e.g., text labels, types) and visual features (e.g., position, size, color).
- Used GPT-4 to evaluate design elements in the JSON representation and generate detailed suggestions for guideline violations.
- Refined feedback into constructive language and presented it to designers via the plugin interface.
- Enabled user interaction (e.g., marking or skipping feedback) to dynamically optimize model outputs based on user input.
- Limited evaluations to single-screen static mockups to accommodate GPT-4's context window constraints.
Research Outcomes
-
Specific Outcomes:
- Developed a scalable Figma plugin for iterative design evaluations.
- Experimental validation showed GPT-4 outperformed other existing language models in heuristic evaluation tasks.
- Proposed direct improvement suggestions for design feedback and tested the practical application of generative AI as a design support tool.
-
Advantages:
- Capable of identifying subtle design errors, improving UI text quality, and optimizing UI semantic organization.
- Fast evaluation speed with detailed and clear suggestions, making it suitable for assisting designers.
- Surpassed individual human evaluators in evaluation depth and detail, particularly in capturing intricate issues.
-
Experimental or Evaluation Results:
- Evaluated 51 UIs, with 52% of suggestions deemed accurate and 49% considered highly helpful.
- Compared with manual evaluations by 12 human experts, GPT-4 identified 9 design issues overlooked by humans.
- Iterative usage studies revealed the model's effectiveness decreased as UIs improved, but its initial evaluation performance was the most impactful.
-
Limitations and Future Directions:
- The model relies on the quality of the UI JSON, particularly the semantic accuracy of group names and labels.
- The current tool supports only single-screen static UIs and cannot evaluate interactive behaviors or multi-screen design consistency.
- The model's understanding of design context is limited, and feedback occasionally contains repetitions or errors.
- Future work could explore leveraging large context windows and multimodal models (e.g., GPT-4V) to handle extensive UIs and improve visual cognition.
- Investigate integrating the tool into real-world design projects, from initial concepts to final implementation.
Overall, this research demonstrates the potential of large language models in design evaluation applications, with the optimization of model capabilities through next-generation AI technologies being a key direction for future development.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- Can large language models such as GPT-4 be used for heuristic automatic evaluation of UI prototypes and generate specific actionable design feedback?Category: UI, 3D, and Visual Design AssistanceSimilar questionsarrow_forward
- How can JSON-represented UI structures support large language models in semantic analysis and hierarchical evaluation of complex interfaces?Category: UI, 3D, and Visual Design AssistanceSimilar questionsarrow_forward
- How can designers dynamically optimize AI-generated feedback results through plugin interaction?Category: UI, 3D, and Visual Design AssistanceSimilar questionsarrow_forward
Practical Problems
1- UI designers struggle to obtain fast and high-quality design evaluation and feedback.Category: UI, 3D, and Visual Design AssistanceSimilar questionsarrow_forward
- 100%
StoryEnsemble: Enabling Dynamic Exploration & Iteration in the Design Process with AI and Forward-Backward Propagation
UIST '25· Human-LLM Collaboration +1
- 80%
May AI? Design Ideation with Cooperative Contextual Bandits
CHI '19· Generative AI (Text, Image, Music, Video) +2
- 80%
Generative AI in User Experience Design and Research: How Do UX Practitioners, Teams, and Companies Use GenAI in Industry?
DIS '24· Generative AI (Text, Image, Music, Video) +2
- 80%
PDFChatAnnotator: A Human-LLM Collaborative Multi-Modal Data Annotation Tool for PDF-Format Catalogs
IUI '24· Human-LLM Collaboration +1
- 75%
Storyboard-Based Empirical Modeling of Touch Interface Performance
CHI '18· Prototyping & User Testing
- 75%
Steering through Successive Objects
CHI '18· Prototyping & User Testing
- 75%
Applied Sketching in HCI: Hands-on Course of Sketching Techniques
CHI '18· Prototyping & User Testing
- 75%
GUIComp: A GUI Design Assistant with Real-Time, Multi-Faceted Feedback
CHI '20· Prototyping & User Testing
- 75%
SimUser: Generating Usability Feedback by Simulating Various Users Interacting with Mobile Applications
CHI '24· Human-LLM Collaboration +1
- 75%
Interaction Substrates: Combining Power and Simplicity in Interactive Systems
CHI '25· Prototyping & User Testing
Based on Jaccard similarity of research subtopics & professions (≥60%)