ILuvUI: Instruction-tuned LangUage-Vision modeling of UIs from Machine Conversations
Multimodal Vision-Language Models (VLMs) enable powerful applications from their fused understanding of images and language, but many perform poorly on UI tasks due to the lack of UI training data. In this paper, we adapt a recipe for generating paired text-image training data for VLMs to the UI domain by combining existing pixel-based methods with a Large Language Model (LLM). Unlike prior art, our method requires no human-provided annotations, and it can be applied to any dataset of UI screenshots. We generate a dataset of 353K conversational examples paired with UIs that cover Q&A, UI descriptions, and planning, and use it to fine-tune a conversational VLM for UI tasks. To assess the performance of our model, we benchmark it on UI element detection tasks, evaluate response quality, and showcase its applicability to UI verification.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- Why do multimodal vision-language models (VLMs) perform poorly on graphical user interface (UI) tasks?Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
- How can high-quality image-text training data for UI tasks be generated without manual annotation?Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
- What performance levels can improved multimodal models achieve in UI element recognition, description generation, and task planning?Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
Practical Problems
1- Users struggle to effectively understand complex interfaces, hindering automation and accessibility technology development.Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
- 100%
Content-Driven Local Response: Supporting Sentence-Level and Message-Level Mobile Email Replies With and Without AI
CHI '25· Voice User Interface (VUI) Design +1
- 100%
Vajra: Step-by-step programming with natural language
IUI '19· Voice User Interface (VUI) Design +1
- 80%
Enabling Conversational Interaction with Mobile UI using Large Language Models
CHI '23· Voice User Interface (VUI) Design +1
- 80%
Screen2Words: Automatic Mobile UI Summarization with Multimodal Learning
UIST '21· Voice User Interface (VUI) Design +1
- 75%
Tap&Say: Touch Location-Informed Large Language Model for Multimodal Text Correction on Smartphones
CHI '25· Human-LLM Collaboration
- 67%
ReactGenie: A Development Framework for Complex Multimodal Interactions Using Large Language Models
CHI '24· Voice User Interface (VUI) Design +2
- 67%
Multimodal Error Correction for Speech-to-Text in a Mobile Office Automated Vehicle: Results From a Remote Study
IUI '22· Automated Driving Interface & Takeover Design +2
- 67%
VoiceAlign: A Shimming Layer for Enhancing the Usability of Legacy Voice User Interface Systems
IUI '26· Voice User Interface (VUI) Design +2
- 67%
ReMap: Lowering the Barrier to Help-Seeking with Multimodal Search
UIST '20· Voice User Interface (VUI) Design +2
- 60%
GestAKey: Touch Interaction on Individual Keycaps
CHI '18· Hand Gesture Recognition +1
Based on Jaccard similarity of research subtopics & professions (≥60%)