Learning Network-Based Multi-Modal Mobile User Interface Embeddings
Title of the Paper
Learning Network-Based Multi-Modal Mobile User Interface Embeddings
Paper Information
- Research Domain: Mobile user interface design, multi-modal network representation learning, deep learning
- Keywords: Network embedding, mobile application user interface, unsupervised retrieval, multi-modal, multi-task learning
Research Background and Problems
-
Identified Problems or Challenges:
- Mobile user interface design encompasses rich multi-modal information (e.g., text, code, images, categories, and numerical data), which current methods fail to effectively capture in a comprehensive semantic manner.
- Existing retrieval systems rely solely on keyword or classification methods, unable to fully leverage the multi-modal and non-Euclidean nature of interface design, leading to inefficiencies in retrieval and recommendation.
- Most embedding generation methods focus on single modalities, failing to effectively capture relationships among multi-modal features within network structures.
-
Importance:
- Cross-modal representation of mobile application user interfaces is crucial for improving retrieval efficiency, accuracy, and user satisfaction in interface design, especially when applied to large-scale real-world datasets like the RICO dataset.
-
Research Motivation and Related Work:
- The authors analyzed the limitations of existing network embedding and user interface retrieval models, including insufficient support for multi-modal information and multi-task learning.
- To address these issues, this paper proposes a novel multi-modal embedding model (MAAN) that aims to comprehensively utilize the multi-modal and heterogeneous network information inherent in user interface design.
Solution
-
Proposed Method or Solution:
- A novel unsupervised model—Multi-modal Attention-Based Attributed Network Embedding (MAAN)—is proposed.
- Based on Graph Variational Autoencoder (GVAE), MAAN effectively integrates multi-modal and network structure information while leveraging an attention mechanism to balance the contributions of different modalities.
-
Innovations:
- The first integration of attention mechanisms with the variational autoencoder framework for multi-task learning (e.g., link prediction, attribute prediction, regression, and retrieval).
- A two-stage encoding process ensures that the generated embeddings are not dominated by any single modality.
- The attention mechanism autonomously discovers the relevance of information across different modalities.
- The introduction of Maximum Mean Discrepancy (MMD) loss into GVAE for the first time to balance multi-task objectives, enhancing training stability and embedding significance.
-
Implementation Steps and Key Techniques:
- Network Structure Representation:
- Construct a heterogeneous bipartite graph consisting of UI screen nodes and UI element nodes, along with edges connecting them.
- Encode node features using multi-head attention mechanisms and generate Gaussian-distributed node embeddings through Graph Convolutional Networks (GCN).
- Multi-modal Information Fusion:
- Create independent Graph Attention Network (GAT) modules for node features of different modalities.
- Adaptively assign weights to features from multiple modalities using attention mechanisms.
- Generating Final Embeddings:
- Compress encoded feature representations into embedding space.
- Reconstruct node and attribute information using an inner product decoder.
- Objective Function Optimization:
- Optimize the graph autoencoder by combining reconstruction loss with KL divergence/MMD loss.
- Network Structure Representation:
Research Outcomes
-
Specific Results:
- MAAN outperformed state-of-the-art models in multiple tasks (e.g., link prediction, UI attribute inference, UI score prediction, UI retrieval).
- Experiments on two RICO datasets demonstrated that MAAN significantly surpassed traditional single-modal embedding methods (e.g., CAN and GAT), particularly in handling continuous-valued attributes and multi-modal retrieval tasks.
-
Advantages Over Existing Solutions:
- MAAN better captures the correlations between modalities and the characteristics of network structures.
- Improvements based on MMD loss enhance the model's training stability and representational capacity.
- User evaluations revealed that MAAN's retrieval performance exceeded existing models by up to 8 times.
-
Experimental and Evaluation Results:
- In the UI screen score prediction task, MAAN achieved the lowest RMSE (0.55), demonstrating accurate modeling of real-world scores.
- For UI screen attribute inference, MAAN significantly improved precision, particularly in predicting continuous attributes (e.g., reducing RMSE for screen images from 2.454 to 0.994).
- In the UI retrieval task, AMT user tests showed that MAAN achieved an average accuracy of 85%, outperforming other methods.
-
Limitations and Future Directions:
- Edge Types: Currently supports only single edge type in structural information; future work could extend to multiple edge types.
- Positional Information: Lacks the ability to capture the sequential hierarchy of UI elements within views, which may affect model performance.
- End-to-End Training: The encoding of certain modalities is not integrated into the main model; future research could develop end-to-end architectures to enhance performance.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can multimodal information and network structure knowledge be used to generate efficient mobile UI embeddings?Category: Control and Co-Creation in Generative CreationSimilar questionsarrow_forward
- Based on attention mechanisms and variational autoencoders, can the stability and representational capacity of multimodal embedding generation be improved?Category: Control and Co-Creation in Generative CreationSimilar questionsarrow_forward
- In multimodal retrieval, how can the weights of information from each modality be better balanced to improve retrieval precision?Category: Control and Co-Creation in Generative CreationSimilar questionsarrow_forward
Practical Problems
1- Designers face inefficiency and imprecision when retrieving mobile app interfaces.Category: Control and Co-Creation in Generative CreationSimilar questionsarrow_forward
- 75%
TiiS: Humanized Recommender Systems: State-of-the-Art and Research Issues
IUI '22· Recommender System UX
- 67%
A decision-theoretic representation of assistive interfaces
CHI '26· AI-Assisted Decision-Making & Automation +2
- 67%
Rethinking User Empowerment in AI Recommender System: Innovating Transparent and Controllable Interfaces
CHI '26· Explainable AI (XAI) +2
- 67%
In-Situ Adaptive Interfaces for Online Browsing: Design Dimensions for Intent-Responsive Automation and User Control
IUI '26· AI-Assisted Decision-Making & Automation +2
- 60%
Designing for the Bittersweet: Improving Sensitive Experiences with Recommender Systems
CHI '22· AI Ethics, Fairness & Accountability +1
- 60%
Predicting and Explaining Mobile UI Tappability with Vision Modeling and Saliency Analysis
CHI '22· Explainable AI (XAI) +1
- 60%
How Experienced Designers of Enterprise Applications Engage AI as a Design Material
CHI '22· Generative AI (Text, Image, Music, Video) +1
- 60%
Two Heads Are Better Than One: A Dimension Space for Unifying Human and Artificial Intelligence in Shared Control
CHI '22· AI-Assisted Decision-Making & Automation +1
- 60%
“Should I Follow the Human, or Follow the Robot?” — Robots in Power Can Have More Influence Than Humans on Decision-Making
CHI '23· AI-Assisted Decision-Making & Automation +1
- 60%
Dealing with Uncertainty: Understanding the Impact of Prognostic Versus Diagnostic Tasks on Trust and Reliance in Human-AI Decision Making
CHI '24· AI-Assisted Decision-Making & Automation +1
Based on Jaccard similarity of research subtopics & professions (≥60%)