GhostUI: Unveiling Hidden Interactions in Mobile UI

Mobile App User ExperienceOne-Handed Operation & Mobile GesturesHuman-LLM CollaborationSoftware Engineers & DevelopersAI/ML Researchers & Engineers

Paper Title

GhostUI: Unveiling Hidden Interactions in Mobile UI

Publication Info

  • Topic area: Detection and modeling of hidden interactions in mobile user interfaces.
  • Keywords: Hidden interactions, mobile UI, vision language models, gesture recognition, dataset, task automation, affordances, multimodal learning, interaction design, mobile agents.

Background and Problem

  • Problem / challenge: Modern mobile UIs increasingly rely on hidden interactions (e.g., long presses, swipes) that lack visual cues, making them difficult for users and vision language models (VLMs) to detect and utilize. Existing datasets focus on visible interactions, leaving a gap in understanding and automating hidden gestures.
  • Significance: Hidden interactions are critical for maximizing functionality in constrained mobile interfaces, but their discoverability and usability remain limited, hindering both user experience and the capabilities of automated mobile agents.
  • Motivation and related work: Prior work on mobile UI datasets and VLMs has focused on visible affordances, neglecting hidden interactions. Current mobile agents lack support for complex gestures like double taps and pinches, limiting their practical deployment. This paper addresses these gaps by systematically documenting hidden interactions and enabling their modeling.

Solution

  • Proposed approach: GhostUI, a dataset and framework specifically designed to document and model hidden interactions in mobile UIs, including paired screenshots, view hierarchies, gesture metadata, and task descriptions.
  • Novelty:
    1. Introduction of GhostUI, a dataset with 1,970 validated hidden interaction instances from 81 popular mobile applications.
    2. Development of a formal taxonomy categorizing six gesture types and their contextual usage patterns.
    3. Demonstration that VLMs fine-tuned on GhostUI outperform baselines in gesture classification and spatial localization tasks.
  • Procedure and key techniques:
    • Automated UI probing tool to systematically test six gesture types across mobile apps.
    • Manual validation of interactions to ensure accuracy and hidden nature.
    • Task contextualization with natural language descriptions for multimodal training.
    • Dataset split into training and testing sets at the app level to ensure generalizability.

Results

  • Concrete findings:
    • GhostUI includes 1,970 hidden interaction instances, with gestures distributed as follows: tap (30.3%), swipe (26.0%), long press (19.3%), double tap (9.5%), pinch (8.9%), and scroll (6.0%).
    • Fine-tuned GPT-4o achieved 65.6% gesture classification accuracy and 42.5% IoU, outperforming its zero-shot baseline by 14.5% and 6.5%, respectively.
    • Fine-tuned Qwen2.5-VL improved to 40.5% accuracy and 22.8% IoU, up from 33.3% and 19.5%.
  • Advantage over baselines:
    • Fine-tuned models showed significant improvements in predicting hidden interactions and UI transitions compared to zero-shot baselines.
    • Ablation studies revealed the critical importance of view hierarchies and gesture usage patterns for spatial localization and classification accuracy.
  • Experiments / evaluation:
    • Two tasks: Hidden Interaction Prediction (gesture classification and localization) and UI Transition Prediction (post-interaction state description).
    • Models evaluated using accuracy, IoU, and cosine similarity metrics.
    • Dataset split ensured no overlap between training and testing apps.
  • Limitations and future work:
    • Limited to Android apps; future work should include iOS and other platforms.
    • Focused on six gestures; additional interaction types like drag-and-drop and sensor-based gestures remain unexplored.
    • Manual annotation is time-intensive; semi-automated methods could improve scalability.

Summary

GhostUI addresses the challenges of modeling hidden interactions in mobile UIs by introducing a comprehensive dataset and framework. It documents 1,970 hidden interaction instances across six gesture types, enabling significant performance improvements in VLMs for gesture classification and UI transition prediction. The dataset highlights the prevalence and complexity of hidden interactions, offering insights for both mobile task automation and interaction design. Future work could expand platform and gesture coverage, improve annotation scalability, and explore human factors to enhance usability and accessibility.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222920/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790283
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Mobile App User Experience, One-Handed Operation & Mobile Gestures, Human-LLM Collaboration
work
Professions
Software Engineers & Developers, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers