Visualization of Speech Prosody and Emotion in Captions: Accessibility for Deaf and Hard-of-Hearing Users
Authors
Title of the Paper
Visualization of Speech Prosody and Emotion in Captions: Accessibility for Deaf and Hard-of-Hearing Users
Paper Information
- Subject Area: Accessibility Technology, Visualization of Speech Emotion and Prosody
- Keywords: Accessibility, Affective Computing, Assistive Technology, Deaf and Hard-of-Hearing Users, Speech Prosody, Visual Text
Research Background and Problem Statement
-
Identified Problems or Challenges:
- Automatically generated captions often fail to convey the emotional and emphasis-related information in speech, making it difficult for Deaf and Hard-of-Hearing (DHH) users to grasp the overall semantics or tone of conversations in online meetings.
- Existing caption designs lack visual representation of speech prosody (volume, pitch, speed) and emotional information, resulting in captions that are overly monotonous and unengaging.
-
Significance:
- Automatic captions have become a vital tool for DHH users in online meetings, but their design limitations restrict users' sense of participation and interaction efficiency, especially in speech exchanges involving emotional and emphasis expressions.
-
Research Motivation and Related Work:
- The study aims to explore how visualizing speech prosody and emotional information can improve the expressive quality of captions, making them more inclusive and effective for DHH users in both professional and personal meetings.
- Previous related studies have rarely addressed the visualization of prosody and emotion in real-time video conferencing, with existing work focusing on pre-recorded content processing or audio modeling.
Proposed Solution
-
Proposed Solution:
- Develop three novel caption models to surpass traditional text captions:
- A model representing speech prosody (including volume, pitch, and duration).
- A model representing speech emotion (based on a two-dimensional emotional model of valence and arousal).
- An integrated model representing both speech prosody and emotion.
- Develop three novel caption models to surpass traditional text captions:
-
Innovative Aspects:
- The new caption models visually express emotional and emphasis-related information by analyzing the acoustic features of speech signals in real time.
- They provide visual markers based on font styles (e.g., font weight, color, baseline shift, letter spacing) to represent prosody and emotion.
-
Implementation Steps and Techniques:
- Step 1: Extract prosodic features (volume, pitch, duration) through automatic speech recognition and acoustic analysis.
- Step 2: Analyze emotional features (valence and arousal) using a Transformer-based neural network model.
- Step 3: Design improved typography for real-time caption display, including dynamic font styles and color mapping, for real-time rendering.
Research Outcomes
-
Specific Outcomes:
- Through two empirical studies:
- The first study interviewed 8 DHH users, identifying the shortcomings of current caption systems and their needs for emotion- and prosody-embedded caption models.
- The second study tested the acceptance and comprehension of different caption models with 16 DHH users, evaluating the effectiveness of the new models in conveying emotional and emphasis-related information.
- Results indicate that the emotion model significantly outperforms traditional captions in expressing emotional and emphasis-related information, while the prosody model is less effective than traditional captions in emphasizing information.
- Through two empirical studies:
-
Advantages Compared to Existing Solutions:
- The new caption models better capture emotional shifts in speech (e.g., transitions from positive to negative), significantly reducing communication ambiguity caused by captions.
- Although slightly less readable than traditional captions, the emotion model excels in conveying narrative emotions, better capturing users' attention.
-
Experimental or Evaluation Results:
- Users rated the emotion captions significantly higher than traditional captions in the dimension of "clearly identifying emotions and tone."
- Additionally, users found the emotion model more suitable for personal meetings, while traditional captions were deemed more appropriate for professional settings.
-
Limitations and Future Directions:
-
Limitations:
- Readability issues in font design: Certain font styles (e.g., baseline shift and letter spacing changes) may reduce caption readability.
- Adaptation for colorblind users: The current color scheme of the caption models may have low distinguishability for colorblind users.
- Insufficient consideration of caption usage in different scenarios (e.g., when facial expressions are not visible).
-
Future Directions:
- Optimize color marking schemes to enhance adaptation for colorblind users.
- Explore more granular speech feature extraction and display (e.g., based on syllables or phonemes).
- Incorporate user- or speaker-controlled emotion input interfaces to improve interactive experiences.
- Further investigate whether the new caption models can enhance interaction quality and immersion in communication.
-
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can visualizing speech prosody (volume, pitch, and speed) improve captioning experience for Deaf and hard-of-hearing users?Category: Captions, Speech Transcription, and Deaf Communication AssistanceSimilar questionsarrow_forward
- How can visualizing emotional information (based on valence and arousal) improve expressiveness of captions in online meetings?Category: Captions, Speech Transcription, and Deaf Communication AssistanceSimilar questionsarrow_forward
- Do combined emotion and prosody caption models outperform traditional captions in expressing emotion and emphasis?Category: Captions, Speech Transcription, and Deaf Communication AssistanceSimilar questionsarrow_forward
Practical Problems
1- Deaf and hard-of-hearing users struggle to obtain emotional and emphasis information from speech through traditional captions.Category: Captions, Speech Transcription, and Deaf Communication AssistanceSimilar questionsarrow_forward
- 100%
Visible Nuances: A Caption System to Visualize Paralinguistic Speech Cues for Deaf and Hard-of-Hearing Individuals
CHI '23· Voice Accessibility +2
- 83%
Watch It, Don't Imagine It: Creating a Better Caption-Occlusion Metric by Collecting More Ecologically Valid Judgments from DHH Viewers
CHI '22· Voice Accessibility +2
- 80%
Automatic Text Simplification Tools for Deaf and Hard of Hearing Adults: Benefits of Lexical Simplification and Providing Users with Autonomy
CHI '20· Voice Accessibility +1
- 67%
Methods for Evaluation of Imperfect Captioning Tools by Deaf or Hard-of-Hearing Users at Different Reading Literacy Levels
CHI '18· Voice Accessibility +1
- 67%
How Users Experience Closed Captions on Live Television: Quality Metrics Remain a Challenge
CHI '24· Voice Accessibility +2
- 67%
Caption Royale: Exploring the Design Space of Affective Captions from the Perspective of Deaf and Hard-of-Hearing Individuals
CHI '24· Voice Accessibility +2
- 67%
Weaving Sound Information to Support Real-Time Sensemaking of Auditory Environments: Co-Designing with a DHH User
CHI '25· Voice Accessibility +2
- 67%
Sounds Accessible: Envisioning Accessible Audio Media Futures with People with Aphasia
CHI '25· Voice Accessibility +2
- 60%
Towards AI-driven Sign Language Generation with Non-manual Markers
CHI '25· Voice Accessibility +1
- 60%
To Each Their Own: Exploring Highly Personalised Audiovisual Media Accessibility Interventions with People with Aphasia
DIS '25· Voice Accessibility +1
Based on Jaccard similarity of research subtopics & professions (≥60%)