Visual Captions: Augmenting Verbal Communication with On-the-fly Visuals

Voice User Interface (VUI) DesignDeaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)Cybersecurity EngineersHCI Researchers

Title of the Paper

Visual Captions: Augmenting Verbal Communication with On-the-fly Visuals

Paper Information

  • Research Area: Human-Computer Interaction and Augmented Communication
  • Keywords: Augmented Communication, Large Language Models, Video Conferencing, Online Meetings, Collaborative Work, Dataset, Text-to-Visual, AI Agents, Augmented Reality

Research Background and Problem Statement

  • Existing Problems or Challenges:
    Current video conferencing platforms like Zoom and Google Meet, despite offering real-time captioning features, have limitations in enhancing the understanding of complex or unfamiliar concepts. Existing text-to-visual conversion technologies are primarily designed for non-real-time text content and cannot address the need for visual augmentation in synchronous human verbal communication. Furthermore, the deployment of real-time systems and their impact on user interaction and communication effectiveness remain underexplored.

  • Significance:
    Integrating visual elements related to conversational semantics in video conferencing can improve communication efficiency, help explain complex or abstract information, and enhance the engagement and enjoyment of conversations. Research has shown that multimodal information is more engaging and effective for comprehension compared to single modalities (e.g., audio only).

  • Motivation and Related Work:
    The authors recognize the necessity of building a synchronous visual augmentation system. Based on the observation that people often use visual materials to supplement verbal communication, this study aims to develop a real-time system that predicts "visual intent" in user conversations and provides relevant visual suggestions.

Solution

  • Proposed Method and Solution:
    The authors developed a real-time system called "Visual Captions," which integrates with video conferencing platforms to provide visual content support for synchronous verbal communication. The system predicts visual intent in user conversations and presents visual content in three different AI proactivity modes.

  • Innovations:

    1. The first system to attempt understanding user visual intent and providing real-time visual content suggestions.
    2. Utilizes a fine-tuned GPT-3 language model to process open-vocabulary conversations and retrieve visual resources from personal photo libraries and online searches.
    3. Offers user-controllable AI proactivity options, including Auto-display, Auto-suggest, and On-demand-suggest modes.
  • Implementation Steps and Key Technologies:

    1. Data Collection and Processing: Built the VC1.5K dataset with 1,595 visual intents by recruiting 246 crowdsourced workers.
    2. Model Training: Fine-tuned the GPT-3 model to predict the type, source, and specific content of visual materials in conversations, achieving a validation accuracy of 86.59%.
    3. User Interface Design: Developed a Chrome browser plugin to support real-time visual augmentation, integrating automatic caption parsing and visual suggestions.
    4. AI Mode Design: The system provides three proactivity modes to cater to different user interaction preferences.

Research Outcomes

  • Specific Achievements:

    1. Dataset Contribution: Released an open dataset containing 1,595 visual intents across 15 conversational topics.
    2. Language Model Optimization: Developed an accurate and robust visual intent prediction language model capable of handling open-vocabulary conversational content.
    3. Visual Captions System: Successfully implemented a real-time enhancement Chrome plugin that enriches communication in video conferencing.
    4. User Testing Results: The provided visual content significantly improved understanding, clarified linguistic ambiguities, and enhanced the overall communication experience.
  • Advantages and Experimental Results:

    1. Over 80% of participants in experiments found the system's visual suggestions helpful for understanding and clarifying concepts.
    2. Users preferred the Auto-suggest mode in one-on-one and casual scenarios, while opting for lower-proactivity modes in formal settings.
    3. Deployment experiments demonstrated that Visual Captions enhanced group interactions and the overall appeal of meetings.
  • Limitations and Future Directions:

    1. Visual Accuracy: Errors in automatic speech recognition may lead to inaccurate visual content. Future improvements could involve corpus-based topic modeling and stability prediction to enhance visual quality.
    2. Model Personalization: Reducing redundant content for familiar concepts and better adapting to individual preferences is necessary.
    3. Diverse Application Scenarios: Exploring system extensions to other domains such as augmented reality, 3D visualization, storytelling, and creative thinking support.
    4. Interaction Mode Optimization: Investigating user adaptation to AI proactivity in different social contexts and designing dynamically adjustable interaction modes.

This study demonstrates the potential of real-time visual augmentation systems to enhance verbal communication and identifies several valuable opportunities for future research directions.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/95817/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581566
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Voice User Interface (VUI) Design, Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)
work
Professions
Cybersecurity Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers