Generating Automatic Feedback on UI Mockups with Large Language Models

Human-LLM CollaborationPrototyping & User TestingUI/UX DesignersHCI Researchers

Title of the Paper

Generating Automatic Feedback on UI Mockups with Large Language Models

Paper Information

  • Research Area: Human-Computer Interaction, Interface Design, Application of AI in Design
  • Keywords: Large Language Models, Human-Computer Interface Design Tools, GPT-4, Heuristic Evaluation, Design Feedback

Research Background and Problem

  • Problem or Challenge: In the UI design process, human feedback is a crucial factor for improving designs. However, experienced evaluators are often difficult to access, and manual feedback is costly. Additionally, existing automated feedback tools are limited in evaluation dimensions and struggle to provide interpretable analysis results.
  • Significance: UI design directly impacts the quality of user interaction with technology and information. Fast and effective feedback can significantly enhance design efficiency while reducing the subjectivity of design quality.
  • Research Motivation: Heuristic evaluation is a widely used design evaluation method but still requires experienced human evaluators to conduct manual checks. The potential of AI-assisted evaluation remains underexplored, particularly for supporting large-scale, automated assessments.
  • Related Work: Existing research leverages computer vision and machine learning techniques to provide design feedback, but retraining models for different tasks has limitations. Emerging applications of generative AI demonstrate new possibilities for design support, but their application to heuristic evaluation of UI remains insufficiently explored.

Solution

  • Method or Solution:

    1. Developed a GPT-4-based large language model tool for automated heuristic evaluation of UI prototypes.
    2. Implemented the tool as a Figma plugin to evaluate static UI mockups and generate constructive textual suggestions.
    3. Introduced a JSON-based representation of UI to enable the model to process complex UI hierarchies and semantic information.
    4. Supported user-defined evaluation guidelines and dynamically adjusted model outputs based on designer feedback.
  • Innovations:

    • Cross-guideline support for diverse design feedback.
    • Emphasis on presenting feedback in context to enhance designers' understanding of its implications.
    • Introduced a reflection mechanism where the model self-corrects based on contextual feedback when designers mark suggestions as incorrect.
  • Implementation Steps and Key Technologies:

    1. Constructed a JSON representation of UI screens, including semantic features (e.g., text labels, types) and visual features (e.g., position, size, color).
    2. Used GPT-4 to evaluate design elements in the JSON representation and generate detailed suggestions for guideline violations.
    3. Refined feedback into constructive language and presented it to designers via the plugin interface.
    4. Enabled user interaction (e.g., marking or skipping feedback) to dynamically optimize model outputs based on user input.
    5. Limited evaluations to single-screen static mockups to accommodate GPT-4's context window constraints.

Research Outcomes

  • Specific Outcomes:

    1. Developed a scalable Figma plugin for iterative design evaluations.
    2. Experimental validation showed GPT-4 outperformed other existing language models in heuristic evaluation tasks.
    3. Proposed direct improvement suggestions for design feedback and tested the practical application of generative AI as a design support tool.
  • Advantages:

    • Capable of identifying subtle design errors, improving UI text quality, and optimizing UI semantic organization.
    • Fast evaluation speed with detailed and clear suggestions, making it suitable for assisting designers.
    • Surpassed individual human evaluators in evaluation depth and detail, particularly in capturing intricate issues.
  • Experimental or Evaluation Results:

    1. Evaluated 51 UIs, with 52% of suggestions deemed accurate and 49% considered highly helpful.
    2. Compared with manual evaluations by 12 human experts, GPT-4 identified 9 design issues overlooked by humans.
    3. Iterative usage studies revealed the model's effectiveness decreased as UIs improved, but its initial evaluation performance was the most impactful.
  • Limitations and Future Directions:

    1. The model relies on the quality of the UI JSON, particularly the semantic accuracy of group names and labels.
    2. The current tool supports only single-screen static UIs and cannot evaluate interactive behaviors or multi-screen design consistency.
    3. The model's understanding of design context is limited, and feedback occasionally contains repetitions or errors.
    4. Future work could explore leveraging large context windows and multimodal models (e.g., GPT-4V) to handle extensive UIs and improve visual cognition.
    5. Investigate integrating the tool into real-world design projects, from initial concepts to final implementation.

Overall, this research demonstrates the potential of large language models in design evaluation applications, with the optimization of model capabilities through next-generation AI technologies being a key direction for future development.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/146712/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642782
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Human-LLM Collaboration, Prototyping & User Testing
work
Professions
UI/UX Designers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers