Human interaction with large language models (LLMs) is typically confined to text or image interfaces. Sketches offer a powerful medium for articulating creative ideas and user intentions, yet their potential remains underexplored. We propose SketchGPT, a novel interaction paradigm that integrates sketch and speech input directly over the system interface, facilitating open-ended, context-aware communication with LLMs. By leveraging the complementary strengths of multimodal inputs, expressions are enriched with semantic scope while maintaining efficiency. Interpreting user intentions across diverse contexts and modalities remains a key challenge. To address this, we developed a prototype based on a multi-agent framework that infers user intentions within context and generates executable context-sensitive and toolkit-aware feedback. Using Chain-of-Thought techniques for temporal and semantic alignment, the system understands multimodal intentions and performs operations following human-in-the-loop confirmation to ensure reliability. User studies demonstrate that SketchGPT significantly outperforms unimodal manipulation approaches, offering more intuitive and effective means to interact with LLMs.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/uist/207018/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3746059.3747598
At a Glance

Paper Snapshot

fact_check
dataset
Source
UIST
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
12 authors
sell
Subtopics
Voice User Interface (VUI) Design, Human-LLM Collaboration
work
Professions
—
article
Content Status
Abstract only
hub
Related Papers
3 related papers