Virtual Minds, Real Work: LLM-Powered Preference-Based Planning through Spatial Multi-Agent-Human Collaboration
Authors
Paper Title
Virtual Minds, Real Work: LLM-Powered Preference-Based Planning through Spatial Multi-Agent-Human Collaboration
Publication Info
- Topic area: Preference-based planning using multi-agent systems and spatial visualization.
- Keywords: Preference-based planning, multi-agent systems, LLMs, spatial visualization, VR, human-AI collaboration, Taskformer, CoNavigator, interactive planning, user-centered design.
Background and Problem
- Problem / challenge: Existing planning systems lack iterative, adaptive interactions, require structured inputs, and fail to effectively integrate nuanced user preferences. Multi-agent LLM systems often operate as closed systems with limited user involvement and transparency.
- Significance: Effective preference-based planning is crucial in domains like travel, project management, and creative tasks, where balancing constraints and evolving user preferences is essential.
- Motivation and related work: Prior systems rely on rigid inputs (e.g., PDDL) and static outputs, limiting their applicability to real-world tasks. LLMs enable natural language interaction but struggle with multidimensional preferences and trade-offs. Multi-agent LLM systems show potential but often exclude users from the planning process. This paper addresses these gaps by designing an interactive, user-centered planning system.
Solution
- Proposed approach: MAVIS (Multi-Agent Virtual Interactive Synergy), a multi-agent system combining incremental collaboration (Taskformer) and spatial visualization (CoNavigator) for preference-based planning.
- Novelty:
- Incremental multi-agent collaboration mechanism (Taskformer) that decomposes tasks into stages and introduces expert agents sequentially.
- Spatial visualization (CoNavigator) externalizing agent reasoning into step-linked summaries and embodied avatars for natural interaction.
- Cross-modality implementation supporting both VR and desktop environments.
- Procedure and key techniques:
- Taskformer stages: Guideline generation, incremental expansion, focused exploration, collaborative negotiation, and consolidation.
- CoNavigator modules: Virtual space for spatially organized content and embodied agents for intuitive interaction.
- Integration of LLMs for role assignment, preference elicitation, and trade-off negotiation.
- Evaluation through three studies comparing MAVIS to baselines and across VR and desktop modalities.
Results
- Concrete findings:
- Taskformer increased preference articulation by 60.3% and improved planning quality significantly over GPT-4o.
- CoNavigator reduced cognitive overload and improved engagement and sensemaking.
- MAVIS supported diverse real-world tasks with high usability ratings (UMUX-Lite: ~81/100).
- Advantage over baselines:
- Higher preference expressed rate (MAS: 0.58 vs. GPT-4o: 0.29, p<0.001).
- Higher preference met rate (MAS: 0.93 vs. GPT-4o: 0.58, p<0.001).
- More balanced time distribution across planning steps and richer multi-agent interactions in MAVIS compared to MAS.
- Experiments / evaluation:
- Study 1: Quantitative comparison of Taskformer vs. GPT-4o on travel planning (N=13).
- Study 2: Qualitative comparison of MAVIS (with CoNavigator) vs. MAS (N=16).
- Study 3: Expert evaluation of MAVIS across VR and desktop modalities in domain-specific tasks (N=14).
- Limitations and future work:
- Reliance on general-purpose LLMs may limit performance in specialized domains.
- Lack of cross-device workflows for transitioning between VR and desktop.
- Fixed agent behaviors and social presence may induce anxiety in some users.
- Limited exploration of richer VR-specific interaction techniques.
Summary
MAVIS introduces a novel approach to preference-based planning by combining incremental multi-agent collaboration (Taskformer) with spatial visualization (CoNavigator). It improves user engagement, preference articulation, and plan alignment compared to traditional LLM systems. MAVIS supports diverse real-world tasks and operates effectively across VR and desktop modalities. Future work should focus on domain-specific enhancements, cross-device workflows, and adaptive social behaviors to further refine multi-agent human-AI collaboration.
Research Questions / Practical Problems
Question signals indexed for this paper.
Based on Jaccard similarity of research subtopics & professions (≥60%)