ReactGenie: A Development Framework for Complex Multimodal Interactions Using Large Language Models

Voice User Interface (VUI) DesignGenerative AI (Text, Image, Music, Video)Human-LLM CollaborationSoftware Engineers & DevelopersUI/UX DesignersAI/ML Researchers & Engineers

Document Title

ReactGenie: A Development Framework for Complex Multimodal Interactions Using Large Language Models

Document Information

  • Subject Area: Human-Computer Interaction, Interface Development, Multimodal Interaction
  • Keywords: Multimodal Interaction, Development Framework, Programming Framework, Large Language Models, Natural Language Processing

Research Background and Problem

  • Identified Issues or Challenges: With the advancement of multimodal interactions (e.g., touch and voice), existing frameworks require developers to manually handle complex commands, leading to high development costs and extended timelines. Additionally, achieving expressiveness and compositionality in language interactions remains challenging.
  • Significance: Multimodal interactions can enhance the efficiency and user experience of interfaces, but the complexity of developing such applications poses a high technical barrier, limiting their widespread adoption.
  • Motivation and Related Work: Existing frameworks (e.g., Redux, QuickSet) primarily handle either voice or graphical interfaces and lack support for the compositional operation of multimodal commands. Therefore, a new framework is needed to simplify the development process and facilitate the implementation of multimodal interactions.

Solution

  • Proposed Method or Solution: The ReactGenie framework redefines the interface layer and computational model by integrating large language models (LLMs). It parses users' multimodal commands into a custom domain-specific programming language (NLPL) and executes these commands through an internal interpreter.
  • Innovations:
    • Introduced object-layered state abstraction to support the natural composition of multimodal commands, reducing intermediate variables.
    • Designed the NLPL programming language and an LLM-based multimodal command parsing model to improve command understanding and semantic parsing accuracy.
    • Automated UI mapping, semantic parsing, and feedback generation, significantly reducing developers' workload.
  • Implementation Steps and Key Technologies:
    1. Developer Programming: Developers define functionalities using object state classes and UI components, marking accessible features via annotations.
    2. Framework Initialization and Module Generation: Automatically generates semantic parsing and UI mapping modules by extracting developer-defined functionalities using LLMs.
    3. Runtime Processing: Listens to user voice and touch operations, transcribes voice input into NLPL code, executes the generated code via an interpreter, updates the UI automatically, and generates textual feedback.

Research Outcomes

  • Specific Results:
    • Developed the ReactGenie framework, providing developers with an easy-to-use tool for building complex multimodal applications.
    • Demonstrated the framework's applicability across three representative domains (e.g., food ordering, social networking, and NDA management).
  • Advantages Compared to Existing Solutions:
    • Significant reduction in development time: Compared to GPT-3 function calls, ReactGenie reduced development time and effort by nearly 75%.
    • High expressiveness: Approximately 90% of user commands were correctly parsed, significantly outperforming traditional semantic parsing frameworks.
  • Experimental or Evaluation Results:
    • Framework Performance Evaluation:
      • 12 developers learned and completed demonstration applications within an average of 2.5 hours.
    • User Interaction Testing:
      • Task completion efficiency improved by approximately 50% (average task time reduced from 63.6 seconds to 33.6 seconds).
      • Cognitive load decreased, and user experience improved, as evaluated using NASA-TLX and SUS standards.
    • Natural language parsing accuracy reached 90%, with no completely erroneous behaviors observed.
  • Limitations and Future Directions:
    1. Voice Interface Improvements: Enhance voice feedback functionality and support more natural multi-turn conversations.
    2. Development Tool Enhancements: Reduce the number of parsing examples required and explore automatic generation of example code.
    3. Support for Additional Modalities: Extend the framework to support complex gestures and other interaction modalities, such as image processing and gesture recognition.

Conclusion

The ReactGenie framework demonstrates a novel approach to supporting complex multimodal interactions through large language models. The research highlights the efficiency, usability, and enhanced user interaction capabilities of ReactGenie. This study provides new insights for future multimodal application development and has the potential to advance the field of human-computer interaction.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147488/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642517
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
10 authors
sell
Subtopics
Voice User Interface (VUI) Design, Generative AI (Text, Image, Music, Video), Human-LLM Collaboration
work
Professions
Software Engineers & Developers, UI/UX Designers, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers