Tactile Emotions: Multimodal Affective Captioning with Haptics Improves Narrative Engagement for d/Deaf and Hard-of-Hearing Viewers

Vibrotactile Feedback & Skin StimulationDeaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)Speech-Language Pathologists & AudiologistsDisability Service Providers

Research Background and Issues

  • What problems or challenges did the authors identify?
    The authors observed limitations in current caption systems designed for deaf and hard-of-hearing individuals, particularly in conveying the speaker's emotions. Existing studies suggest that visual cues can effectively communicate emotional valence (the positivity or negativity of emotions) but are less effective in conveying emotional arousal (intensity). Additionally, relying solely on visual cues may not fully meet user needs, as users with hearing impairments require a more comprehensive perception of emotional information to enhance their video-watching experience.

  • Why is this issue important?
    For deaf and hard-of-hearing users, captions are a core tool for understanding speech content and serve as a primary means for learning and entertainment. Improving the ability of captions to convey non-verbal information, such as emotions, can significantly enhance the narrative experience and overall engagement of these users.

  • Research Motivation and Related Work
    Building on previous findings that font color can effectively represent emotional valence but fails to clearly convey emotional arousal, the authors hypothesize that relying solely on visual methods may not be ideal for communicating emotional intensity. They propose combining haptic feedback (via wearable devices) to enhance the transmission of emotional information. This cross-modal design aims to better meet the diverse needs of users.


Solution

  • What methods or solutions did the authors propose?
    The authors designed an innovative multimodal caption system that combines visual and haptic cues to convey the speaker's emotions. Specifically, the system uses visual methods to communicate emotional valence and haptic vibrations (provided by a wrist-worn device) to convey emotional arousal.

  • What are the innovative aspects of this solution?

    • Introducing haptic feedback as a supplementary channel for conveying emotional intensity, a novel approach in caption design.
    • Conducting detailed evaluations of vibration patterns and frequencies to optimize the effectiveness and comfort of the haptic design.
    • Integrating haptic and visual cues to demonstrate the potential of cross-modal information transmission.
  • What are the implementation steps and key technologies used?

    1. Caption Generation: Using open-source tools (e.g., OpenAI's Whisper model) to generate time-stamped captions and adding metadata for emotional valence and arousal to each word via emotion recognition tools.
    2. Haptic Feedback Design: Designing six haptic patterns combining different rhythms (continuous vibration, single short pulse, multiple short pulses) and frequencies (75Hz and 250Hz), and selecting user-preferred settings through experiments.
    3. Visual Cue Design: Adjusting the visual style of captions, including font color (to convey valence) and font weight (to convey arousal).
    4. System Implementation: Using scripts written in the ChucK programming language to generate haptic signals corresponding to textual emotions and synchronizing them with wearable devices.
    5. User Testing: Conducting two studies to evaluate user preferences and narrative engagement for each mode.

Research Outcomes

  • What specific outcomes were achieved?

    • Study 1: Identified the user-preferred haptic pattern as the low-frequency (75Hz) single short pulse mode. This setting effectively conveyed emotional intensity while maintaining user comfort.
    • Study 2: The combination of visual and haptic cues in captions significantly enhanced narrative engagement for deaf and hard-of-hearing users, outperforming traditional plain captions and captions with only visual emotional cues.
  • What advantages does it have compared to existing solutions?

    • The cross-modal design offers significant advantages in conveying complex emotions, reducing information loss that may occur with single-channel visual or auditory methods.
    • The user-preferred haptic pattern minimizes common visual distraction issues while enhancing the transmission of emotional changes.
    • Experimental results demonstrate that captions incorporating haptic feedback improve users' emotional connection and immersion with the narrative content.
  • What were the experimental or evaluation results?
    The research found:

    • The combined visual and haptic caption mode (c4V+H) achieved the highest scores in "narrative comprehension," "attention focus," "narrative presence," and "emotional engagement," significantly outperforming traditional captions (c1B) and captions with only visual emotional cues (c2V).
    • User feedback indicated that this combined mode more clearly conveyed emotional changes, enhancing the overall video experience.
  • Limitations and Future Directions

    • Limitations:
      • Prolonged use of haptic feedback systems may cause discomfort or distraction for some users, requiring further research into user sensitivity and adaptability.
      • The current design is limited to laboratory settings and needs validation in real-world scenarios (e.g., on mobile phones or TVs).
      • The real-time performance and accuracy of the emotion recognition model in captions may not have been fully addressed.
    • Future Directions:
      • Exploring the integration of haptic signals with sound to achieve richer cross-modal expression.
      • Considering personalized user needs to improve design flexibility.
      • Adding adaptive adjustment mechanisms to reduce haptic feedback interference for users.

Overall, this paper proposes an innovative multimodal caption system design that integrates visual and haptic cues to significantly enhance deaf and hard-of-hearing users' comprehension and engagement with video content. It has the potential for widespread application in assistive technology fields.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188499/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713304
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Vibrotactile Feedback & Skin Stimulation, Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)
work
Professions
Speech-Language Pathologists & Audiologists, Disability Service Providers
article
Content Status
Full text indexed
hub
Related Papers
2 related papers