SketchGPT: A Sketch-based Multimodal Interface for Application-Agnostic LLM Interaction
Authors
Human interaction with large language models (LLMs) is typically confined to text or image interfaces. Sketches offer a powerful medium for articulating creative ideas and user intentions, yet their potential remains underexplored. We propose SketchGPT, a novel interaction paradigm that integrates sketch and speech input directly over the system interface, facilitating open-ended, context-aware communication with LLMs. By leveraging the complementary strengths of multimodal inputs, expressions are enriched with semantic scope while maintaining efficiency. Interpreting user intentions across diverse contexts and modalities remains a key challenge. To address this, we developed a prototype based on a multi-agent framework that infers user intentions within context and generates executable context-sensitive and toolkit-aware feedback. Using Chain-of-Thought techniques for temporal and semantic alignment, the system understands multimodal intentions and performs operations following human-in-the-loop confirmation to ensure reliability. User studies demonstrate that SketchGPT significantly outperforms unimodal manipulation approaches, offering more intuitive and effective means to interact with LLMs.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
Automatically Generating and Improving Voice Command Interface from Operation Sequences on Smartphones
CHI '22· Voice User Interface (VUI) Design +1
- 67%
Multi-Modal Approaches for Post-Editing Machine Translation
CHI '19· Voice User Interface (VUI) Design +1
- 67%
Rambler: Supporting Writing With Speech via LLM-Assisted Gist Manipulation
CHI '24· Voice User Interface (VUI) Design +2
Based on Jaccard similarity of research subtopics & professions (≥60%)