ILuvUI: Instruction-tuned LangUage-Vision modeling of UIs from Machine Conversations

Voice User Interface (VUI) DesignHuman-LLM CollaborationSoftware Engineers & DevelopersUI/UX Designers

Multimodal Vision-Language Models (VLMs) enable powerful applications from their fused understanding of images and language, but many perform poorly on UI tasks due to the lack of UI training data. In this paper, we adapt a recipe for generating paired text-image training data for VLMs to the UI domain by combining existing pixel-based methods with a Large Language Model (LLM). Unlike prior art, our method requires no human-provided annotations, and it can be applied to any dataset of UI screenshots. We generate a dataset of 353K conversational examples paired with UIs that cover Q&A, UI descriptions, and planning, and use it to fine-tune a conversational VLM for UI tasks. To assess the performance of our model, we benchmark it on UI element detection tasks, evaluate response quality, and showcase its applicability to UI verification.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/iui/195813/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3708359.3712129
At a Glance

Paper Snapshot

fact_check
dataset
Source
IUI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Voice User Interface (VUI) Design, Human-LLM Collaboration
work
Professions
Software Engineers & Developers, UI/UX Designers
article
Content Status
Abstract only
hub
Related Papers
10 related papers