GenComUI: Exploring Generative Visual Aids as Medium to Support Task-Oriented Human-Robot Communication
Authors
Research Background and Issues
-
Issues and Challenges:
- Current natural language-based robot task programming faces an abstraction matching problem: the gap between natural language expressions and actual program code makes it difficult for users to define complex tasks.
- Sole reliance on voice interaction may lead to unclear expressions of complex logic or spatial task requirements and is constrained by speech recognition errors.
- Traditional rule-based or template-based feedback systems struggle to dynamically adapt to users' evolving needs.
-
Significance:
- With the widespread application of service robots in fields such as home, education, healthcare, and retail, providing intuitive and effective task interfaces for non-technical users is becoming increasingly important.
- Dynamic visual media (e.g., diagrams, path indications, animations) may become crucial tools for enhancing human-computer interaction and user programming experiences.
-
Research Motivation and Related Work:
- The combination of dynamic user interfaces and artificial intelligence (especially large language models, LLMs) holds the potential for automatically generating visual aids, enhancing multimodal interaction through voice and visuals.
- Drawing inspiration from the extensive use of visual aids in interpersonal communication (e.g., charts, arrows) can help convey information more clearly. However, there is currently a lack of research on the application of dynamically generated visual aids in robot interaction.
Solution
-
Proposed Solution:
- A system named "GenComUI" is proposed, which dynamically generates visual aids driven by LLMs, combining voice interaction for task customization and confirmation.
- Integrates a multimodal feedback mechanism that combines graphical animations with voice output to confirm and refine robot task requirements.
-
Innovations:
- Dynamically generated visual aids are highly synchronized with natural language interactions, enabling users to intuitively understand and adjust tasks.
- Based on interaction context, visual feedback is used to display changes in real time, with animations highlighting added or removed task steps.
- Explores how humans use visual tools in task communication to design natural and adaptive human-computer interaction models.
-
Implementation Steps and Key Technologies:
- Modular Architecture Design:
- Voice Interaction Module: Supports speech recognition and conversion for bidirectional communication.
- User Intent Understanding Module: Analyzes user input based on context, generates task steps, and provides feedback.
- Generative Visual Aid Module: Generates dynamic graphical interface elements based on dialogue context.
- Task Program Synthesis and Deployment Module: Produces executable robot operation code based on task instructions.
- Key Technologies:
- Uses LLMs (e.g., GPT-4) for intent parsing, generating structured outputs in real time (including task steps, voice feedback, and visual designs).
- Dynamically constructs task flows using animations and graphical elements based on predefined APIs.
- Generates iterative visual task representations during multi-turn dialogues for step-by-step confirmation and optimization.
- Modular Architecture Design:
Research Outcomes
-
Specific Achievements:
- System Design and Implementation: Developed GenComUI, capable of dynamically generating multimodal feedback to support complex task voice programming.
- User Study: Validated the significant support provided by visual aids in task communication through comparisons with baseline systems.
- Design Guidelines: Summarized strategies for adapting and applying visual aids in complex task communication.
-
Advantages Over Existing Solutions:
- The dynamic, context-aware visual generation mechanism is more aligned with user needs compared to traditional static feedback.
- Helps users track the progress of task communication in real time, performing exceptionally well in complex spatial and logical task processing.
- User testing indicates the system significantly reduces cognitive load during communication and enhances the accuracy and intuitiveness of task development.
-
Experimental or Evaluation Results:
- User Feedback: GenComUI demonstrates significant advantages in helping users express intentions, confirm task understanding, and remember task steps.
- Quantitative Results:
- Compared to baseline systems, GenComUI offers greater interaction flexibility in task confirmation and modification processes.
- Chatbot Usability Scale and System Usability Scale scores show higher user approval for GenComUI (p<0.01).
- Observed Effects:
- In complex task setups, users can more quickly track the robot's understanding of tasks and immediately correct errors.
- Users are more inclined to choose GenComUI (19/20), particularly when handling complex logical branches and large-scale spatial tasks.
-
Limitations and Future Directions:
- Limitations:
- The user group primarily consists of university students and staff, limiting representativeness.
- The experimental environment is a simulated office, which may not encompass the diversity of real-world task scenarios.
- Current tasks are mainly focused on spatial navigation, requiring further research on other task types (e.g., time scheduling).
- Future Directions:
- Extend system functionality to real-world applications, supporting diverse tasks and visual representations.
- Explore the real-time generation of more complex visual elements and study their applicability in different task communications.
- Deepen adaptability to user preferences and task complexity, offering more freedom and personalized options.
- Limitations:
This research opens new directions in the integration of dynamic visual aids and voice interaction, providing significant insights into the combination of human-computer interaction and natural language programming. These findings lay the foundation for designing more intuitive and natural human-computer interaction interfaces in the future.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can dynamically generated visual aids improve UX in service robot task programming?Category: Scientific, Cultural, and Domain Data AnalyticsSimilar questionsarrow_forward
- How can dynamic visual feedback enhance user understanding and confirmation of complex task steps in multi-turn dialogue?Category: Scientific, Cultural, and Domain Data AnalyticsSimilar questionsarrow_forward
- Can LLM-driven visualization feedback significantly reduce users' cognitive load in task definition processes?Category: Scientific, Cultural, and Domain Data AnalyticsSimilar questionsarrow_forward
Practical Problems
1- Non-technical users struggle to accurately define complex tasks using natural language.Category: Scientific, Cultural, and Domain Data AnalyticsSimilar questionsarrow_forward
- 100%
New Enactions of Expertise: Software Engineers’ Evaluation and Demonstration of Coding Expertise with AI Coding Assistants
CHI '26· Human-LLM Collaboration +2
- 100%
From Junior to Senior: Allocating Agency and Navigating Professional Growth in Agentic AI-Mediated Software Engineering
CHI '26· Human-LLM Collaboration +2
- 100%
Developer Interaction Patterns with Proactive AI: A Five-Day Field Study
IUI '26· AI-Assisted Decision-Making & Automation +2
- 83%
Vibe Coding Entanglements – Repositioning Boundaries of Intention, Authorship, and Responsibility in Programming with Generative AI
CHI '26· Generative AI (Text, Image, Music, Video) +2
- 83%
Gemini at Work: Knowledge Workers' Perceptions and Assessment of Productivity Gains
DIS '25· Generative AI (Text, Image, Music, Video) +3
- 83%
PREDILECT: Preferences Delineated with Zero-Shot Language-based Reasoning in Reinforcement Learning
HRI '24· Generative AI (Text, Image, Music, Video) +2
- 83%
GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and Learning
UIST '24· Generative AI (Text, Image, Music, Video) +2
- 80%
Discovering the Syntax and Strategies of Natural Language Programming with Generative Language Models
CHI '22· Generative AI (Text, Image, Music, Video) +1
- 80%
Comparing Sentence-Level Suggestions to Message-Level Suggestions in AI-Mediated Communication
CHI '23· Human-LLM Collaboration +1
- 80%
"What It Wants Me To Say": Bridging the Abstraction Gap Between End-User Programmers and Code-Generating Large Language Models
CHI '23· Generative AI (Text, Image, Music, Video) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)