GenComUI: Exploring Generative Visual Aids as Medium to Support Task-Oriented Human-Robot Communication

Generative AI (Text, Image, Music, Video)Human-LLM CollaborationAI-Assisted Decision-Making & AutomationSoftware Engineers & DevelopersAI/ML Researchers & Engineers

Research Background and Issues

  • Issues and Challenges:

    • Current natural language-based robot task programming faces an abstraction matching problem: the gap between natural language expressions and actual program code makes it difficult for users to define complex tasks.
    • Sole reliance on voice interaction may lead to unclear expressions of complex logic or spatial task requirements and is constrained by speech recognition errors.
    • Traditional rule-based or template-based feedback systems struggle to dynamically adapt to users' evolving needs.
  • Significance:

    • With the widespread application of service robots in fields such as home, education, healthcare, and retail, providing intuitive and effective task interfaces for non-technical users is becoming increasingly important.
    • Dynamic visual media (e.g., diagrams, path indications, animations) may become crucial tools for enhancing human-computer interaction and user programming experiences.
  • Research Motivation and Related Work:

    • The combination of dynamic user interfaces and artificial intelligence (especially large language models, LLMs) holds the potential for automatically generating visual aids, enhancing multimodal interaction through voice and visuals.
    • Drawing inspiration from the extensive use of visual aids in interpersonal communication (e.g., charts, arrows) can help convey information more clearly. However, there is currently a lack of research on the application of dynamically generated visual aids in robot interaction.

Solution

  • Proposed Solution:

    • A system named "GenComUI" is proposed, which dynamically generates visual aids driven by LLMs, combining voice interaction for task customization and confirmation.
    • Integrates a multimodal feedback mechanism that combines graphical animations with voice output to confirm and refine robot task requirements.
  • Innovations:

    • Dynamically generated visual aids are highly synchronized with natural language interactions, enabling users to intuitively understand and adjust tasks.
    • Based on interaction context, visual feedback is used to display changes in real time, with animations highlighting added or removed task steps.
    • Explores how humans use visual tools in task communication to design natural and adaptive human-computer interaction models.
  • Implementation Steps and Key Technologies:

    • Modular Architecture Design:
      1. Voice Interaction Module: Supports speech recognition and conversion for bidirectional communication.
      2. User Intent Understanding Module: Analyzes user input based on context, generates task steps, and provides feedback.
      3. Generative Visual Aid Module: Generates dynamic graphical interface elements based on dialogue context.
      4. Task Program Synthesis and Deployment Module: Produces executable robot operation code based on task instructions.
    • Key Technologies:
      • Uses LLMs (e.g., GPT-4) for intent parsing, generating structured outputs in real time (including task steps, voice feedback, and visual designs).
      • Dynamically constructs task flows using animations and graphical elements based on predefined APIs.
      • Generates iterative visual task representations during multi-turn dialogues for step-by-step confirmation and optimization.

Research Outcomes

  • Specific Achievements:

    1. System Design and Implementation: Developed GenComUI, capable of dynamically generating multimodal feedback to support complex task voice programming.
    2. User Study: Validated the significant support provided by visual aids in task communication through comparisons with baseline systems.
    3. Design Guidelines: Summarized strategies for adapting and applying visual aids in complex task communication.
  • Advantages Over Existing Solutions:

    • The dynamic, context-aware visual generation mechanism is more aligned with user needs compared to traditional static feedback.
    • Helps users track the progress of task communication in real time, performing exceptionally well in complex spatial and logical task processing.
    • User testing indicates the system significantly reduces cognitive load during communication and enhances the accuracy and intuitiveness of task development.
  • Experimental or Evaluation Results:

    • User Feedback: GenComUI demonstrates significant advantages in helping users express intentions, confirm task understanding, and remember task steps.
    • Quantitative Results:
      • Compared to baseline systems, GenComUI offers greater interaction flexibility in task confirmation and modification processes.
      • Chatbot Usability Scale and System Usability Scale scores show higher user approval for GenComUI (p<0.01).
    • Observed Effects:
      • In complex task setups, users can more quickly track the robot's understanding of tasks and immediately correct errors.
      • Users are more inclined to choose GenComUI (19/20), particularly when handling complex logical branches and large-scale spatial tasks.
  • Limitations and Future Directions:

    • Limitations:
      • The user group primarily consists of university students and staff, limiting representativeness.
      • The experimental environment is a simulated office, which may not encompass the diversity of real-world task scenarios.
      • Current tasks are mainly focused on spatial navigation, requiring further research on other task types (e.g., time scheduling).
    • Future Directions:
      1. Extend system functionality to real-world applications, supporting diverse tasks and visual representations.
      2. Explore the real-time generation of more complex visual elements and study their applicability in different task communications.
      3. Deepen adaptability to user preferences and task complexity, offering more freedom and personalized options.

This research opens new directions in the integration of dynamic visual aids and voice interaction, providing significant insights into the combination of human-computer interaction and natural language programming. These findings lay the foundation for designing more intuitive and natural human-computer interaction interfaces in the future.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/189092/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3714238
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Human-LLM Collaboration, AI-Assisted Decision-Making & Automation
work
Professions
Software Engineers & Developers, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers