Winder: Linking Speech and Visual Objects to Support Communication in Asynchronous Collaboration
Authors
Voice User Interface (VUI) DesignMixed Reality WorkspacesCreative Collaboration & Feedback SystemsVocational Trainers & CoachesGame Developers & DesignersUI/UX Designers
Document Title
Winder: Linking Speech and Visual Objects to Support Communication in Asynchronous Collaboration
Document Information
- Topic Area: Design and evaluation of multimodal communication tools in asynchronous collaboration
- Keywords: asynchronous collaboration, teamwork, speech interaction, multimodal input, visual documents, user interface design, information navigation
Research Background and Problem Statement
-
Problem Statement:
- Teams engaged in asynchronous collaboration often experience communication delays or burdens when sharing visual documents (e.g., UI designs or presentations). These issues include high physical and cognitive costs of text input and ambiguous references to visual objects, leading to misunderstandings in collaboration.
- Team members may fail to respond promptly or miss critical information, while text-based communication modes cannot fully meet the needs of multimodal interaction.
-
Importance:
- Efficient communication is crucial for establishing shared understanding and enhancing team productivity. For student teams, it aids in learning collaboration techniques and prevents communication breakdowns.
- Literature highlights that information delays in asynchronous communication can harm productivity and even affect social cognition (e.g., trust among team members).
-
Research Motivation:
- Asynchronous collaboration requires lighter communication compared to face-to-face (F2F) interactions, while still effectively conveying intentions and avoiding misunderstandings.
- Current tools (e.g., Google Docs commenting features) fail to alleviate these burdens, necessitating a tool that integrates speech and object references efficiently.
Solution
-
Proposed Method and Tool:
- Winder is a Figma-based plugin designed to support multimodal communication in asynchronous collaboration. Its core feature, "linked tapes," integrates speech and click operations.
- The tool creates bidirectional links between visual objects in documents and voice recordings, enabling users to quickly understand context through interactive methods.
-
Innovations:
- Bidirectional Linking: Establish connections between document objects and voice recordings. Users can filter relevant recordings by clicking objects or highlight specific document objects during voice playback.
- Pre-recording Feature: Users can record information during current tasks for future reference, mitigating delays in asynchronous communication.
- Multimodal Support: Combines voice comments and click operations to facilitate quick expression and precise object referencing.
-
Implementation Steps and Key Technologies:
- Creation of Linked Tapes: Users record audio in real-time while clicking on relevant objects in the design document (highlighting the objects).
- Enhanced Understanding and Navigation:
- (a) Automatic object highlighting during voice playback;
- (b) Embedding object thumbnails in voice transcripts;
- (c) Object-based recording search.
- Technical Support:
- Using the Figma API to retrieve object thumbnails;
- Google Cloud Speech-to-Text service for voice transcription;
- Frontend built with TypeScript and ReactJS, backend with Node.js and MongoDB.
Research Outcomes
-
Achievements:
- Users experienced lower burdens when generating linked tapes, with interaction processes resembling F2F communication.
- Bidirectional linking improved the accuracy of communication content (object and voice information) and reduced redundant work.
- Participants recorded an average of 13.13 tapes, with an average length of 53.27 seconds, staying within acceptable user limits.
-
Comparison with Existing Solutions:
- Compared to traditional text-based asynchronous communication tools, Winder more accurately conveys intentions and context.
- ITS features (e.g., object highlighting, search, thumbnails) not only reduce the sender's burden but also alleviate the receiver's workload.
-
Experimental or Evaluation Results:
- Quantitative Analysis: Among 107 recorded tapes, 83.8% were used for descriptions, 46.6% for explaining design motivations, and 40% for task coordination.
- User Feedback:
- Both click and voice input simplified interaction processes; voice input was easier than typing but still posed challenges in pronunciation or clear expression.
- Object-based search functionality improved information retrieval efficiency and helped avoid team conflicts.
- Social Effects: Pre-recording enhanced trust, motivation, and confidence among team members.
-
Limitations and Future Directions:
- Limitations:
- The study fixed team sizes at three members, which may limit applicability in broader scenarios.
- Participants were non-professional UI designers, so performance in real work environments might differ.
- Future Directions:
- Explore more flexible trigger mechanisms to replace the fixed 10-minute prompts;
- Automatically generate advanced reflective recording guidance;
- Expand application scenarios (e.g., presentations or video discussions).
- Limitations:
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- In asynchronous collaboration, how can combining voice and visual objects reduce communication misunderstandings and information delay?Category: Display Layout, Visual Load, and Presentation PerceptionSimilar questionsarrow_forward
- Can multimodal support (e.g., voice and click operations) improve asynchronous teams' efficiency in understanding design intent and context?Category: Display Layout, Visual Load, and Presentation PerceptionSimilar questionsarrow_forward
- How do bidirectional linking features (associations between voice recordings and document objects) affect communication quality and user burden in asynchronous collaboration?Category: Display Layout, Visual Load, and Presentation PerceptionSimilar questionsarrow_forward
lightbulb
Practical Problems
1- In team asynchronous collaboration, text communication easily creates high burden and reference ambiguity.Category: Display Layout, Visual Load, and Presentation PerceptionSimilar questionsarrow_forward
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3411764.3445686
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2021
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Voice User Interface (VUI) Design, Mixed Reality Workspaces, Creative Collaboration & Feedback Systems
work
Professions
Vocational Trainers & Coaches, Game Developers & Designers, UI/UX Designers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers