Caption Royale: Exploring the Design Space of Affective Captions from the Perspective of Deaf and Hard-of-Hearing Individuals
Honorable MentionAuthors
Voice AccessibilityDeaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)Universal & Inclusive DesignSpeech-Language Pathologists & AudiologistsAssistive Technology Specialists
Title of the Paper
Caption Royale: Exploring the Design Space of Affective Captions from the Perspective of Deaf and Hard-of-Hearing Individuals
Paper Information
- Research Area: Accessibility Design and Human-Computer Interaction (HCI)
- Keywords: Accessibility, Caption Design, Emotional Expression, Deaf and Hard-of-Hearing, Human-Computer Interaction, Typography Design, Emotion Recognition Models, Diverse Experimental Methods
Research Background and Problem Statement
-
Research Questions or Challenges:
- Current captioning technologies have made significant progress in accurately conveying linguistic content but primarily focus on verbal information, lacking the expression of emotion or tone.
- Deaf and Hard-of-Hearing (DHH) audiences miss out on emotional cues conveyed through non-verbal vocal information, impacting their content experience.
- Existing studies have preliminarily explored conveying emotions through caption styles (e.g., using color, font adjustments, etc.), but there has been no systematic investigation of DHH individuals' preferences and the effectiveness of these design styles.
-
Significance:
- Providing accurate and emotionally rich captions for DHH individuals can significantly enhance their video content experience.
- Systematic exploration of emotional caption design can not only enrich design approaches in the HCI field but also advance accessibility technologies.
-
Research Motivation:
- Limited existing work has preliminarily shown that "emotional captions" can enhance the understanding of emotional contexts for DHH individuals, but broad empirical validation from the user perspective is lacking.
- The importance of scrolling captions (word-by-word display) for emotional descriptions among deaf users has not been thoroughly studied, representing a significant gap.
Proposed Solution
-
Methods and Solutions:
- Propose nine distinct caption styles that use attributes such as font color, font weight, size, and shadow to express two dimensions of emotion: valence (positive or negative emotions) and arousal (level of excitement).
- Employ a three-phase research design:
- Study 1: Evaluate the performance of these nine styles in conveying valence and arousal independently.
- Study 2: Combine the best-performing styles from Study 1 to explore their effectiveness in conveying both valence and arousal together.
- Study 3: Compare the optimal styles identified in the first two studies with neutral styles in terms of subjective and objective effects.
-
Innovations:
- Conduct the first systematic and comprehensive analysis of the visual design space for "emotional captions."
- Combine user preferences with emotional conveyance effectiveness, directly comparing which styles better serve DHH individuals.
- Utilize unique experimental methods (e.g., EmojiGrid for two-dimensional emotion rating, Best-Worst Scaling) to validate style effectiveness from both preference and emotional interpretation perspectives.
-
Implementation Steps and Key Techniques:
- Data Preparation: Extract transcripts, timestamps, and vocal emotions (based on emotion recognition networks) from emotionally rich video materials.
- Style Design: Select font parameters most relevant to emotions from the literature and design experimental styles.
- Data Collection: Recruit DHH participants to conduct multi-round tests evaluating style preferences and interpretation effectiveness.
- Data Analysis: Quantify style performance using ranking algorithms (e.g., TrueSkill) and correlation tests (e.g., distance correlation).
Research Outcomes
-
Specific Findings:
- Study 1 Results:
- For valence: Font color and shadow color were the most effective.
- For arousal: Font color, font weight, font size, and shadow color performed well.
- Study 2 Results:
- Optimal combinations included: font color with font weight, font color with font size, font color with shadow color, and shadow color with font size.
- Study 3 Results:
- Font color combined with font weight and font color combined with font size were the most effective, demonstrating distinct advantages in cognitive load and emotional conveyance efficiency.
- Study 1 Results:
-
Advantages:
- Performance: Subjective preference rankings and results from emotion recognition tasks showed that the above two combination styles significantly outperformed traditional neutral captions.
- Flexibility: Participant feedback supported further personalization through parameter adjustments (e.g., color intensity, size, weight range).
-
Experimental and Evaluation Results:
- Subjective Perception: Participants found the combination of font color and font size to be the most intuitive for conveying emotions, while the font color and font weight combination offered lower visual interference and cognitive load.
- Objective Testing: The two optimal styles (Font-color with Font-size and Font-color with Font-weight) showed significantly higher correlation with reference emotional coordinates in emotion recognition tasks compared to other styles.
- Limitations:
- Emotional captions may introduce visual fatigue for prolonged video viewing.
- The cultural differences in color-emotion associations require further validation.
-
Future Directions:
- Test additional conversational scenarios (e.g., multi-speaker dialogues, real-time captions).
- Investigate the impact of cultural differences on the interpretation of emotional captions.
- Optimize the real-time processing capabilities of emotional caption systems for dynamic and natural environments.
- Explore broader user personalization options to enhance flexibility.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How do existing subtitle designs express valence and arousal in emotional dimensions?Category: Health, Emotion, and Supportive Interaction DesignSimilar questionsarrow_forward
- Which subtitle design combinations best enhance emotional experience for deaf and hard-of-hearing users?Category: Health, Emotion, and Supportive Interaction DesignSimilar questionsarrow_forward
- How can deaf and hard-of-hearing users' preferences and effects of emotional subtitles be systematically evaluated?Category: Health, Emotion, and Supportive Interaction DesignSimilar questionsarrow_forward
lightbulb
Practical Problems
1- Deaf and hard-of-hearing viewers struggle to obtain emotional cues from existing subtitles.Category: Health, Emotion, and Supportive Interaction DesignSimilar questionsarrow_forward
- 100%
Weaving Sound Information to Support Real-Time Sensemaking of Auditory Environments: Co-Designing with a DHH User
CHI '25· Voice Accessibility +2
- 67%
Methods for Evaluation of Imperfect Captioning Tools by Deaf or Hard-of-Hearing Users at Different Reading Literacy Levels
CHI '18· Voice Accessibility +1
- 67%
Methods for Evaluating the Fluency of Automatically Simplified Texts with Deaf and Hard-of-Hearing Adults at Various Literacy Levels
CHI '22· Voice Accessibility +2
- 67%
Visualization of Speech Prosody and Emotion in Captions: Accessibility for Deaf and Hard-of-Hearing Users
CHI '23· Voice Accessibility +2
- 67%
Visible Nuances: A Caption System to Visualize Paralinguistic Speech Cues for Deaf and Hard-of-Hearing Individuals
CHI '23· Voice Accessibility +2
- 67%
How Users Experience Closed Captions on Live Television: Quality Metrics Remain a Challenge
CHI '24· Voice Accessibility +2
- 67%
A Sound Understanding --- An In-Situ Deployment of an Accessible Audio-Media Player with People Living with Aphasia
CHI '26· Voice Accessibility +2
- 67%
Vid2Coach: Transforming How-To Videos into Task Assistants
UIST '25· Voice Accessibility +2
- 60%
Towards AI-driven Sign Language Generation with Non-manual Markers
CHI '25· Voice Accessibility +1
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642258
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
Honorable Mention
group
Authors
5 authors
sell
Subtopics
Voice Accessibility, Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration), Universal & Inclusive Design
work
Professions
Speech-Language Pathologists & Audiologists, Assistive Technology Specialists
article
Content Status
Full text indexed
hub
Related Papers
9 related papers