Leveraging Multimodal LLM for Inspirational User Interface Search
Authors
Research Background and Issues
What problems or challenges did the authors identify?
The authors highlighted the following key problems and challenges in current methods for inspiration search in user interface (UI) design:
- Limitations of existing search methods:
- Retrieval methods based on pixel similarity (e.g., platforms like Pinterest, Behance, and Dribbble) primarily focus on visual styles and lack deep semantic support, such as identifying target users, application functional domains, or emotional tones.
- Metadata-dependent methods (e.g., view hierarchy) are constrained by the quality and availability of metadata, limiting their scalability.
- Manual annotation burden on existing platforms:
- For instance, UI curation platforms like Mobbin rely on manual classification of designs, inherently limiting their scalability and comprehensiveness.
- Lack of comprehensive understanding of designers' needs:
- Designers have more granular semantic needs, such as functionality, user interaction flows, target audiences, and emotional tones, which existing tools fail to adequately support.
Why is this issue important?
Inspirational search is a critical step in the UI design process, as designers rely on searching existing interface designs to spark creativity. However, current methods fail to meet designers' needs for functional and visual semantic searches, thereby hindering their creative expression and design efficiency.
Research Motivation and Related Work
- Research Motivation: To improve UI design inspiration search by going beyond visual elements and integrating semantic-level information (e.g., target users, application roles, and emotional tones), making search results more aligned with designers' needs.
- Related Work: Current methods leveraging deep learning and large models offer new possibilities, such as using the CLIP model for visual-semantic relevance retrieval in UIs. However, these methods still face limitations in handling complex semantics, particularly in contextual understanding and interaction analysis.
Solution
What methods or solutions did the authors propose?
-
Using Multimodal Large Language Models (MLLM) to process UI semantic information:
- Leveraging MLLM to directly extract UI semantic information from screenshots without relying on metadata or manual annotations.
- Proposed 23 semantic features, including application category, target users, screen roles, interface layout, and visual design semantics such as emotions and colors.
-
Semantic-driven UI Retrieval System (S&UI):
- Designed and implemented a retrieval system, S&UI, that supports filtering by multiple semantic categories, enabling designers to easily find results using high-precision semantic tags.
- Introduced a weight adjustment feature, allowing users to dynamically prioritize query semantics based on search preferences.
What are the innovative aspects of this solution?
- Direct processing of UI screenshots: Utilizes MLLM to automatically generate new structured semantic data from UI visual information, instead of relying on existing metadata.
- Integration of functionality and aesthetics: Combines functional (e.g., screen roles) and visual semantics (e.g., emotions, colors) in a single system to support comprehensive search.
- Semantic transparency and interpretability: The retrieval system provides detailed semantic explanations, offering contextual support for search results and helping designers quickly understand and evaluate sources of inspiration.
What are the implementation steps? What key technologies were used?
- Formative Research: Conducted interviews with designers to identify the importance of UI semantic elements, categorized into four levels: application, screen, component, and visual design semantics.
- Semantic Information Extraction:
- Used GPT-4o multimodal large language model for semantic analysis of UI screenshots.
- Defined and structured semantic categories (e.g., screen types, emotions, colors) in YAML format to ensure clarity and ease of parsing.
- Development of the Semantic Retrieval System:
- Built the S&UI system to support queries based on semantic descriptions (e.g., "welcome screens for health apps").
- Added a "weight adjustment" feature, enabling designers to dynamically prioritize search parameters.
- User Testing and Experimental Validation:
- Combined manual and computational evaluations to validate the accuracy of semantic extraction, relevance of query results, and system usability.
Research Outcomes
What specific results were achieved?
-
Accurate Semantic Extraction Capability:
- For complex semantics such as application categories and screen roles, GPT-4o significantly outperformed existing CLIP models, achieving a Top-1 classification accuracy of 59.21%.
- Provided unexpected (serendipitous) semantic information, such as target users and emotional tones, broadening users' design perspectives.
-
Superiority in Semantic Retrieval:
- S&UI significantly improved retrieval quality, with users noting clear advantages in query relevance, diversity, and result reliability. Compared to baseline systems (e.g., GUIClip), S&UI demonstrated significant improvements across various metrics (e.g., relevance increased from 4.9 in GUIClip to 6.7 in S&UI).
-
Alignment with Human Evaluation:
- Feedback from human designers indicated that the functionalities, aesthetics, and iterative support provided by S&UI greatly enhanced overall design efficiency.
What advantages does it have compared to existing solutions?
- Comparison with traditional tools:
- Provides UI examples that better match semantics compared to visually focused platforms like Pinterest or Behance.
- Offers greater dynamic interaction capabilities than UI curation tools like Mobbin, particularly for customized searches involving emotions, roles, and target users.
- Strong scalability: Does not require manual annotation or complex metadata, making it applicable to diverse UI datasets.
What were the experimental or evaluation results?
- Tests on mobile UI datasets (e.g., Enrico and CLAY) demonstrated high accuracy and comprehensiveness in semantic information extraction across most categories.
- User task evaluations showed that S&UI provided designers with unexpectedly diverse inspiration, with high scores concentrated on the novelty and depth of semantics.
Limitations and Future Directions
-
Limitations:
- Current semantic queries rely on manual input, which may not be user-friendly for designers in the exploratory phase.
- Retrieval methods may experience decreased relevance as query complexity increases.
- Semantic extraction faces challenges in scenarios with high layout and UI element complexity.
-
Future Directions:
- Introduce open-ended semantic query functionality to enable more dynamic searches by automatically parsing free-text inputs.
- Optimize visual design and layout understanding by integrating more specialized visual language models or traditional deep learning methods.
- Expand to multi-screen and cross-application analysis to further support comprehensive user flow design exploration.
Conclusion
This study successfully developed a semantic-driven UI design search tool by integrating multimodal large language models, providing designers with support in functionality, aesthetics, and semantic transparency. This establishes a foundation for the next generation of intelligent and practical UI inspiration search tools, with the potential to significantly enhance future design systems.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can multimodal large language models (e.g., GPT-4o) directly extract semantic information from UI designs?Category: UI, 3D, and Visual Design AssistanceSimilar questionsarrow_forward
- How effective is the semantic-driven UI retrieval system (e.g., S&UI) at meeting designers' functional and emotional search needs?Category: UI, 3D, and Visual Design AssistanceSimilar questionsarrow_forward
- Where do existing methods fall short compared with design search tools integrating visual and semantic information?Category: UI, 3D, and Visual Design AssistanceSimilar questionsarrow_forward
Practical Problems
1- Designers lack tools supporting functional and emotional semantics in UI inspiration search.Category: UI, 3D, and Visual Design AssistanceSimilar questionsarrow_forward
- 67%
Design and Analysis of Intelligent Text Entry Systems with Function Structure Models and Envelope Analysis
CHI '21· Human-LLM Collaboration +1
- 67%
A Pragmatics-based Approach to Proactive Digital Assistants for Data Exploration
CUI '25· Human-LLM Collaboration +1
- 60%
A Human-Computer Collaborative Editing Tool for Conceptual Diagrams
CHI '23· Human-LLM Collaboration +1
- 60%
WaitGPT: Monitoring and Steering Conversational LLM Agent in Data Analysis with On-the-Fly Code Visualization
UIST '24· Human-LLM Collaboration +1
Based on Jaccard similarity of research subtopics & professions (≥60%)