Visible Nuances: A Caption System to Visualize Paralinguistic Speech Cues for Deaf and Hard-of-Hearing Individuals
Authors
Title of the Paper
Visible Nuances: A Caption System to Visualize Paralinguistic Speech Cues for Deaf and Hard-of-Hearing Individuals
Paper Information
- Subject Area: Accessibility technology, specifically video caption systems designed for individuals with hearing impairments
- Keywords: Caption design, speech accessibility, paralinguistic cues, deaf individuals, emotion visualization, visualization, audio visualization
Research Background and Problem
- Current caption systems primarily convey textual content, neglecting paralinguistic speech cues (such as intonation, tone, and intensity), which prevents individuals with hearing impairments from fully understanding the speaker's intent and emotions.
- Existing improvements (e.g., manually inputting emotional information or displaying basic audio features through specific methods) have limitations: manual input is time-consuming and labor-intensive, and the display methods are insufficient for expressing complex speech emotions.
- Conveying paralinguistic speech cues through captions for individuals with hearing impairments is an important but unresolved research challenge.
Solution
- Overall Approach: Propose an audio visualization caption system that transforms paralinguistic speech cues (volume, pitch, emotion, etc.) into visual elements (such as font thickness, text height, font style, and motion).
- Innovations:
- For the first time, dynamic fonts, text motion, and other advanced visual elements are applied in captions to convey complex speech emotions and intonation.
- The system integrates audio analysis algorithms with predefined visual mapping rules.
- Implementation Steps:
- Audio Separation: Use the audio source separation model (Spleeter) to extract speech from video.
- Feature Extraction and Emotion Classification: Segment audio samples into syllable units, extract volume and pitch features, and use an emotion classifier to identify emotions.
- Caption Generation: Map font style, motion, thickness, and height of captions based on audio features and emotion results to generate real-time dynamic captions.
- Key Technologies:
- Audio separation techniques
- Volume and pitch feature normalization algorithms
- Design for classification and mapping of complex emotions
Research Outcomes
- Specific Results:
- Developed a dynamic caption prototype system for individuals with hearing impairments, capable of visually presenting paralinguistic speech details in videos.
- Validated the system's practicality and user experience with 20 hearing-impaired participants.
- Advantages Compared to Existing Solutions:
- More intuitively conveys emotions and tone, enhancing speech accessibility in video content.
- Effectively helps hearing-impaired individuals recognize the speaker's tone and intent in formal speech videos.
- Experiment or Evaluation Results:
- In formal speech videos, participants using the advanced caption system improved their accuracy in identifying the speaker's intent from 22.5% to 45%.
- In emotion recognition tasks, the advanced caption system increased the accuracy of hearing-impaired individuals in recognizing nuanced emotional subcategories (e.g., sarcasm, loneliness) by 20%.
- The advanced caption system demonstrated significant advantages in content expressiveness and user satisfaction, especially for speech videos where tone is difficult to interpret through facial expressions.
- Limitations and Future Directions:
- Limitations:
- Some participants found the dynamic caption system visually distracting, such as reading fatigue caused by text motion.
- The current emotion classification system has limited coverage, supporting only six emotion categories.
- The system relies on manual emotion annotation and has not yet achieved full automation.
- Future Directions:
- Optimize caption design to reduce visual burden (e.g., refine font styles and motion frequency).
- Expand the system's emotion classification capabilities to support more nuanced emotional expressions.
- Enhance system automation by developing real-time emotion classification and audio analysis modules.
- Explore the natural alignment between caption systems and video content to avoid disconnection between captions and video context.
- Limitations:
Conclusion
This study investigates an innovative caption system for individuals with hearing impairments, capable of visually representing paralinguistic speech information through detailed visual elements. Although the system requires further refinement in visual comfort and design precision, it significantly improves the perception and user experience of speech accessibility in video content for hearing-impaired individuals, providing valuable insights for future caption system design.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can existing captioning systems be improved to convey paralinguistic information (e.g., tone and emotion) to people with hearing loss?Category: Deaf and Hard-of-Hearing ASL and Caption SupportSimilar questionsarrow_forward
- Can visual elements (e.g., font weight, text height, motion) effectively represent complex speech emotional information?Category: Deaf and Hard-of-Hearing ASL and Caption SupportSimilar questionsarrow_forward
- How effective are dynamic captioning systems at enhancing deaf and hard-of-hearing users' understanding of video content?Category: Deaf and Hard-of-Hearing ASL and Caption SupportSimilar questionsarrow_forward
Practical Problems
1- Deaf and hard-of-hearing people struggle to understand speech intonation and emotion through existing captioning systems.Category: Deaf and Hard-of-Hearing ASL and Caption SupportSimilar questionsarrow_forward
- 100%
Visualization of Speech Prosody and Emotion in Captions: Accessibility for Deaf and Hard-of-Hearing Users
CHI '23· Voice Accessibility +2
- 83%
Watch It, Don't Imagine It: Creating a Better Caption-Occlusion Metric by Collecting More Ecologically Valid Judgments from DHH Viewers
CHI '22· Voice Accessibility +2
- 80%
Automatic Text Simplification Tools for Deaf and Hard of Hearing Adults: Benefits of Lexical Simplification and Providing Users with Autonomy
CHI '20· Voice Accessibility +1
- 67%
Methods for Evaluation of Imperfect Captioning Tools by Deaf or Hard-of-Hearing Users at Different Reading Literacy Levels
CHI '18· Voice Accessibility +1
- 67%
How Users Experience Closed Captions on Live Television: Quality Metrics Remain a Challenge
CHI '24· Voice Accessibility +2
- 67%
Caption Royale: Exploring the Design Space of Affective Captions from the Perspective of Deaf and Hard-of-Hearing Individuals
CHI '24· Voice Accessibility +2
- 67%
Weaving Sound Information to Support Real-Time Sensemaking of Auditory Environments: Co-Designing with a DHH User
CHI '25· Voice Accessibility +2
- 67%
Sounds Accessible: Envisioning Accessible Audio Media Futures with People with Aphasia
CHI '25· Voice Accessibility +2
- 60%
Towards AI-driven Sign Language Generation with Non-manual Markers
CHI '25· Voice Accessibility +1
- 60%
To Each Their Own: Exploring Highly Personalised Audiovisual Media Accessibility Interventions with People with Aphasia
DIS '25· Voice Accessibility +1
Based on Jaccard similarity of research subtopics & professions (≥60%)