ReactGenie: A Development Framework for Complex Multimodal Interactions Using Large Language Models
Authors
Karina Li
Computer ScienceYuhan Zhang
Electrical EngineeringDocument Title
ReactGenie: A Development Framework for Complex Multimodal Interactions Using Large Language Models
Document Information
- Subject Area: Human-Computer Interaction, Interface Development, Multimodal Interaction
- Keywords: Multimodal Interaction, Development Framework, Programming Framework, Large Language Models, Natural Language Processing
Research Background and Problem
- Identified Issues or Challenges: With the advancement of multimodal interactions (e.g., touch and voice), existing frameworks require developers to manually handle complex commands, leading to high development costs and extended timelines. Additionally, achieving expressiveness and compositionality in language interactions remains challenging.
- Significance: Multimodal interactions can enhance the efficiency and user experience of interfaces, but the complexity of developing such applications poses a high technical barrier, limiting their widespread adoption.
- Motivation and Related Work: Existing frameworks (e.g., Redux, QuickSet) primarily handle either voice or graphical interfaces and lack support for the compositional operation of multimodal commands. Therefore, a new framework is needed to simplify the development process and facilitate the implementation of multimodal interactions.
Solution
- Proposed Method or Solution: The ReactGenie framework redefines the interface layer and computational model by integrating large language models (LLMs). It parses users' multimodal commands into a custom domain-specific programming language (NLPL) and executes these commands through an internal interpreter.
- Innovations:
- Introduced object-layered state abstraction to support the natural composition of multimodal commands, reducing intermediate variables.
- Designed the NLPL programming language and an LLM-based multimodal command parsing model to improve command understanding and semantic parsing accuracy.
- Automated UI mapping, semantic parsing, and feedback generation, significantly reducing developers' workload.
- Implementation Steps and Key Technologies:
- Developer Programming: Developers define functionalities using object state classes and UI components, marking accessible features via annotations.
- Framework Initialization and Module Generation: Automatically generates semantic parsing and UI mapping modules by extracting developer-defined functionalities using LLMs.
- Runtime Processing: Listens to user voice and touch operations, transcribes voice input into NLPL code, executes the generated code via an interpreter, updates the UI automatically, and generates textual feedback.
Research Outcomes
- Specific Results:
- Developed the ReactGenie framework, providing developers with an easy-to-use tool for building complex multimodal applications.
- Demonstrated the framework's applicability across three representative domains (e.g., food ordering, social networking, and NDA management).
- Advantages Compared to Existing Solutions:
- Significant reduction in development time: Compared to GPT-3 function calls, ReactGenie reduced development time and effort by nearly 75%.
- High expressiveness: Approximately 90% of user commands were correctly parsed, significantly outperforming traditional semantic parsing frameworks.
- Experimental or Evaluation Results:
- Framework Performance Evaluation:
- 12 developers learned and completed demonstration applications within an average of 2.5 hours.
- User Interaction Testing:
- Task completion efficiency improved by approximately 50% (average task time reduced from 63.6 seconds to 33.6 seconds).
- Cognitive load decreased, and user experience improved, as evaluated using NASA-TLX and SUS standards.
- Natural language parsing accuracy reached 90%, with no completely erroneous behaviors observed.
- Framework Performance Evaluation:
- Limitations and Future Directions:
- Voice Interface Improvements: Enhance voice feedback functionality and support more natural multi-turn conversations.
- Development Tool Enhancements: Reduce the number of parsing examples required and explore automatic generation of example code.
- Support for Additional Modalities: Extend the framework to support complex gestures and other interaction modalities, such as image processing and gesture recognition.
Conclusion
The ReactGenie framework demonstrates a novel approach to supporting complex multimodal interactions through large language models. The research highlights the efficiency, usability, and enhanced user interaction capabilities of ReactGenie. This study provides new insights for future multimodal application development and has the potential to advance the field of human-computer interaction.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can large language models simplify development workflows and lower technical barriers in multimodal interaction?Category: GUI/IoT Task Automation and Interface GenerationSimilar questionsarrow_forward
- Can large language models improve semantic parsing accuracy of user commands through custom programming languages?Category: GUI/IoT Task Automation and Interface GenerationSimilar questionsarrow_forward
- How does the ReactGenie framework improve development efficiency and user experience of multimodal interaction applications?Category: GUI/IoT Task Automation and Interface GenerationSimilar questionsarrow_forward
Practical Problems
1- Developers struggle to efficiently implement complex multimodal interaction combining speech and touch.Category: GUI/IoT Task Automation and Interface GenerationSimilar questionsarrow_forward
- 83%
Enabling Conversational Interaction with Mobile UI using Large Language Models
CHI '23· Voice User Interface (VUI) Design +1
- 83%
MUD: Towards a Large-Scale and Noise-Filtered UI Dataset for Modern Style UI Modeling
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 83%
How the Role of Generative AI Shapes Perceptions of Value in Human-AI Collaborative Work
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 83%
Screen2Words: Automatic Mobile UI Summarization with Multimodal Learning
UIST '21· Voice User Interface (VUI) Design +1
- 71%
AI-Augmented Brainwriting: Investigating the use of LLMs in group ideation
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 71%
Generative AI in Knowledge Work: Design Implications for Data Navigation and Decision-Making
CHI '25· Generative AI (Text, Image, Music, Video) +2
- 71%
Beyond Code Generation: LLM-supported Exploration of the Program Design Space
CHI '25· Generative AI (Text, Image, Music, Video) +2
- 71%
CARING-AI: Towards Authoring Context-aware Augmented Reality INstruction through Generative Artificial Intelligence
CHI '25· AR Navigation & Context Awareness +2
- 71%
Prototyping with Prompts: Emerging Approaches and Challenges in Generative AI Design for Collaborative Software Teams
CHI '25· Generative AI (Text, Image, Music, Video) +2
- 71%
Creating Design Resources to Scaffold the Ideation of AI Concepts
DIS '23· Generative AI (Text, Image, Music, Video) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)