Visualization of Speech Prosody and Emotion in Captions: Accessibility for Deaf and Hard-of-Hearing Users

Voice AccessibilityDeaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)Universal & Inclusive DesignSpeech-Language Pathologists & AudiologistsDisability Service Providers

Title of the Paper

Visualization of Speech Prosody and Emotion in Captions: Accessibility for Deaf and Hard-of-Hearing Users

Paper Information

  • Subject Area: Accessibility Technology, Visualization of Speech Emotion and Prosody
  • Keywords: Accessibility, Affective Computing, Assistive Technology, Deaf and Hard-of-Hearing Users, Speech Prosody, Visual Text

Research Background and Problem Statement

  • Identified Problems or Challenges:

    • Automatically generated captions often fail to convey the emotional and emphasis-related information in speech, making it difficult for Deaf and Hard-of-Hearing (DHH) users to grasp the overall semantics or tone of conversations in online meetings.
    • Existing caption designs lack visual representation of speech prosody (volume, pitch, speed) and emotional information, resulting in captions that are overly monotonous and unengaging.
  • Significance:

    • Automatic captions have become a vital tool for DHH users in online meetings, but their design limitations restrict users' sense of participation and interaction efficiency, especially in speech exchanges involving emotional and emphasis expressions.
  • Research Motivation and Related Work:

    • The study aims to explore how visualizing speech prosody and emotional information can improve the expressive quality of captions, making them more inclusive and effective for DHH users in both professional and personal meetings.
    • Previous related studies have rarely addressed the visualization of prosody and emotion in real-time video conferencing, with existing work focusing on pre-recorded content processing or audio modeling.

Proposed Solution

  • Proposed Solution:

    • Develop three novel caption models to surpass traditional text captions:
      1. A model representing speech prosody (including volume, pitch, and duration).
      2. A model representing speech emotion (based on a two-dimensional emotional model of valence and arousal).
      3. An integrated model representing both speech prosody and emotion.
  • Innovative Aspects:

    • The new caption models visually express emotional and emphasis-related information by analyzing the acoustic features of speech signals in real time.
    • They provide visual markers based on font styles (e.g., font weight, color, baseline shift, letter spacing) to represent prosody and emotion.
  • Implementation Steps and Techniques:

    • Step 1: Extract prosodic features (volume, pitch, duration) through automatic speech recognition and acoustic analysis.
    • Step 2: Analyze emotional features (valence and arousal) using a Transformer-based neural network model.
    • Step 3: Design improved typography for real-time caption display, including dynamic font styles and color mapping, for real-time rendering.

Research Outcomes

  • Specific Outcomes:

    • Through two empirical studies:
      1. The first study interviewed 8 DHH users, identifying the shortcomings of current caption systems and their needs for emotion- and prosody-embedded caption models.
      2. The second study tested the acceptance and comprehension of different caption models with 16 DHH users, evaluating the effectiveness of the new models in conveying emotional and emphasis-related information.
    • Results indicate that the emotion model significantly outperforms traditional captions in expressing emotional and emphasis-related information, while the prosody model is less effective than traditional captions in emphasizing information.
  • Advantages Compared to Existing Solutions:

    • The new caption models better capture emotional shifts in speech (e.g., transitions from positive to negative), significantly reducing communication ambiguity caused by captions.
    • Although slightly less readable than traditional captions, the emotion model excels in conveying narrative emotions, better capturing users' attention.
  • Experimental or Evaluation Results:

    • Users rated the emotion captions significantly higher than traditional captions in the dimension of "clearly identifying emotions and tone."
    • Additionally, users found the emotion model more suitable for personal meetings, while traditional captions were deemed more appropriate for professional settings.
  • Limitations and Future Directions:

    • Limitations:

      • Readability issues in font design: Certain font styles (e.g., baseline shift and letter spacing changes) may reduce caption readability.
      • Adaptation for colorblind users: The current color scheme of the caption models may have low distinguishability for colorblind users.
      • Insufficient consideration of caption usage in different scenarios (e.g., when facial expressions are not visible).
    • Future Directions:

      • Optimize color marking schemes to enhance adaptation for colorblind users.
      • Explore more granular speech feature extraction and display (e.g., based on syllables or phonemes).
      • Incorporate user- or speaker-controlled emotion input interfaces to improve interactive experiences.
      • Further investigate whether the new caption models can enhance interaction quality and immersion in communication.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/95801/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581511
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Voice Accessibility, Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration), Universal & Inclusive Design
work
Professions
Speech-Language Pathologists & Audiologists, Disability Service Providers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers