Gesturing Toward Abstraction: Multimodal Convention Formation in Collaborative Physical Tasks
Authors
Paper Title
Gesturing Toward Abstraction: Multimodal Convention Formation in Collaborative Physical Tasks
Publication Info
- Topic area: Multimodal communication and convention formation in collaborative physical tasks.
- Keywords: multimodal communication, convention formation, augmented reality, gesture, speech, collaborative assembly, abstraction, Rational Speech Act, computational modeling, human-computer interaction.
Background and Problem
- Problem / challenge: Prior studies have focused on unimodal or 2D collaborative tasks, leaving a gap in understanding how multimodal communication evolves during repeated physical collaboration. There is limited work on modeling multimodal conventions in iterative, multi-step physical tasks.
- Significance: Understanding how multimodal communication evolves is crucial for designing intelligent agents capable of forming conventions with humans for efficient collaboration in physical environments.
- Motivation and related work: Previous research has shown that people form ad hoc conventions over repeated interactions, using concise and abstract expressions. However, most studies have been unimodal, and the dynamics of multimodal communication (speech and gestures) in physical tasks remain underexplored. This paper builds on prior work by examining how multimodal signals evolve and modeling these changes computationally.
Solution
- Proposed approach: The study investigates how linguistic and gestural conventions emerge in collaborative physical tasks through two studies (unimodal online and multimodal AR-based) and introduces a computational model extending the Rational Speech Act (RSA) framework to multimodal settings.
- Novelty:
- Conducted a large-scale online unimodal study and a controlled AR-mediated multimodal lab study.
- Developed a multimodal dataset for repeated physical collaboration tasks.
- Proposed a computational model capturing abstraction formation and modality preferences in multimodal communication.
- Procedure and key techniques:
- Unimodal study: Participants coordinated on linguistic abstractions in a block assembly task, shifting from block-level to tower-level descriptions over repetitions.
- Multimodal study: Participants used speech and gestures in an AR-mediated physical assembly task, forming conventions and adapting modality use over time.
- Computational model: Extended RSA to multimodal settings, incorporating abstraction mapping ambiguity, utterance ambiguity, and modality preferences. Simulated abstraction formation and modality shifts observed in the studies.
Results
- Concrete findings:
- In the unimodal study, reconstruction accuracy improved from F1 = 0.88 to F1 = 0.98, and instruction length decreased by 50% across repetitions.
- In the multimodal study, task success improved from 74.03% to 98.36%, and instruction and construction times significantly decreased across repetitions.
- Participants introduced tower-level abstractions early and used redundancy in speech and gestures to emphasize changes in position and orientation.
- Advantage over baselines:
- Demonstrated how multimodal communication (speech and gestures) evolves more effectively than unimodal communication, with redundancy enhancing comprehension and efficiency.
- The computational model successfully simulated abstraction formation and diverging modality preferences, aligning with empirical findings.
- Experiments / evaluation:
- Unimodal study: 146 participants (73 dyads) performed a block assembly task over 12 trials, with linguistic abstraction and efficiency analyzed.
- Multimodal study: 40 participants (20 dyads) performed a physical assembly task using AR, with speech and gestures analyzed for abstraction, redundancy, and modality shifts.
- Metrics included task accuracy, instruction length, abstraction level, and modality use.
- Limitations and future work:
- Limited to simple block towers; future work should explore more complex physical structures (e.g., furniture assembly).
- Turn-based protocol limits real-time dynamics; future studies should examine synchronous communication.
- Computational model assumptions (e.g., uniform priors, perfect alignment) require refinement for real-world applications.
Summary
This paper investigates how multimodal communication (speech and gestures) evolves during repeated collaborative physical tasks. Through an online unimodal study and an AR-mediated multimodal study, it demonstrates that participants form linguistic and gestural conventions, shifting from block-level to tower-level abstractions and using redundancy to emphasize changes. A computational model extending the Rational Speech Act framework captures these behaviors, simulating abstraction formation and modality preferences. These findings inform the design of convention-aware intelligent agents for physical assembly tasks, capable of learning and adapting to users’ multimodal abstractions and communication preferences over time.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 63%
PeriphAR: Fast and Accurate Real-World Object Selection with Peripheral Augmented Reality Displays
CHI '26· AR Navigation & Context Awareness +2
- 63%
Can We Infer Object Pose Changes from Hand Movements?
CHI '26· Hand Gesture Recognition +2
Based on Jaccard similarity of research subtopics & professions (≥60%)