Winder: Linking Speech and Visual Objects to Support Communication in Asynchronous Collaboration

Voice User Interface (VUI) DesignMixed Reality WorkspacesCreative Collaboration & Feedback SystemsVocational Trainers & CoachesGame Developers & DesignersUI/UX Designers

Document Title

Winder: Linking Speech and Visual Objects to Support Communication in Asynchronous Collaboration

Document Information

  • Topic Area: Design and evaluation of multimodal communication tools in asynchronous collaboration
  • Keywords: asynchronous collaboration, teamwork, speech interaction, multimodal input, visual documents, user interface design, information navigation

Research Background and Problem Statement

  • Problem Statement:

    • Teams engaged in asynchronous collaboration often experience communication delays or burdens when sharing visual documents (e.g., UI designs or presentations). These issues include high physical and cognitive costs of text input and ambiguous references to visual objects, leading to misunderstandings in collaboration.
    • Team members may fail to respond promptly or miss critical information, while text-based communication modes cannot fully meet the needs of multimodal interaction.
  • Importance:

    • Efficient communication is crucial for establishing shared understanding and enhancing team productivity. For student teams, it aids in learning collaboration techniques and prevents communication breakdowns.
    • Literature highlights that information delays in asynchronous communication can harm productivity and even affect social cognition (e.g., trust among team members).
  • Research Motivation:

    • Asynchronous collaboration requires lighter communication compared to face-to-face (F2F) interactions, while still effectively conveying intentions and avoiding misunderstandings.
    • Current tools (e.g., Google Docs commenting features) fail to alleviate these burdens, necessitating a tool that integrates speech and object references efficiently.

Solution

  • Proposed Method and Tool:

    • Winder is a Figma-based plugin designed to support multimodal communication in asynchronous collaboration. Its core feature, "linked tapes," integrates speech and click operations.
    • The tool creates bidirectional links between visual objects in documents and voice recordings, enabling users to quickly understand context through interactive methods.
  • Innovations:

    1. Bidirectional Linking: Establish connections between document objects and voice recordings. Users can filter relevant recordings by clicking objects or highlight specific document objects during voice playback.
    2. Pre-recording Feature: Users can record information during current tasks for future reference, mitigating delays in asynchronous communication.
    3. Multimodal Support: Combines voice comments and click operations to facilitate quick expression and precise object referencing.
  • Implementation Steps and Key Technologies:

    1. Creation of Linked Tapes: Users record audio in real-time while clicking on relevant objects in the design document (highlighting the objects).
    2. Enhanced Understanding and Navigation:
      • (a) Automatic object highlighting during voice playback;
      • (b) Embedding object thumbnails in voice transcripts;
      • (c) Object-based recording search.
    3. Technical Support:
      • Using the Figma API to retrieve object thumbnails;
      • Google Cloud Speech-to-Text service for voice transcription;
      • Frontend built with TypeScript and ReactJS, backend with Node.js and MongoDB.

Research Outcomes

  • Achievements:

    • Users experienced lower burdens when generating linked tapes, with interaction processes resembling F2F communication.
    • Bidirectional linking improved the accuracy of communication content (object and voice information) and reduced redundant work.
    • Participants recorded an average of 13.13 tapes, with an average length of 53.27 seconds, staying within acceptable user limits.
  • Comparison with Existing Solutions:

    • Compared to traditional text-based asynchronous communication tools, Winder more accurately conveys intentions and context.
    • ITS features (e.g., object highlighting, search, thumbnails) not only reduce the sender's burden but also alleviate the receiver's workload.
  • Experimental or Evaluation Results:

    • Quantitative Analysis: Among 107 recorded tapes, 83.8% were used for descriptions, 46.6% for explaining design motivations, and 40% for task coordination.
    • User Feedback:
      • Both click and voice input simplified interaction processes; voice input was easier than typing but still posed challenges in pronunciation or clear expression.
      • Object-based search functionality improved information retrieval efficiency and helped avoid team conflicts.
    • Social Effects: Pre-recording enhanced trust, motivation, and confidence among team members.
  • Limitations and Future Directions:

    • Limitations:
      1. The study fixed team sizes at three members, which may limit applicability in broader scenarios.
      2. Participants were non-professional UI designers, so performance in real work environments might differ.
    • Future Directions:
      1. Explore more flexible trigger mechanisms to replace the fixed 10-minute prompts;
      2. Automatically generate advanced reflective recording guidance;
      3. Expand application scenarios (e.g., presentations or video discussions).

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/47712/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3411764.3445686
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2021
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Voice User Interface (VUI) Design, Mixed Reality Workspaces, Creative Collaboration & Feedback Systems
work
Professions
Vocational Trainers & Coaches, Game Developers & Designers, UI/UX Designers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers