Haptic-Captioning: Using Audio-Haptic Interfaces to Enhance Speaker Indication in Real Time Captions for Deaf and Hard-of-Hearing Viewers
Authors
Title of the Paper
Haptic-Captioning: Using Audio-Haptic Interfaces to Enhance Speaker Indication in Real-Time Captions for Deaf and Hard-of-Hearing Viewers
Paper Information
- Subject Area: Haptic assistive technology and captioning research, focusing on speaker indication in real-time captions for deaf and hard-of-hearing (DHH) users.
- Keywords: Haptic feedback, captions, accessibility technology, deaf and hard-of-hearing users, multimodal feedback
Research Background and Problem
-
What problems or challenges did the authors identify?
- In real-time captioning scenarios, DHH users face difficulties in identifying speaker switches, especially in captions generated through live transcription or automatic speech recognition.
- Existing visually enhanced caption designs (e.g., color coding, avatar indicators) may suffer from information delays, visual fatigue, and other cognitive burdens in real-time environments.
-
Why is this problem important?
- Captions are a crucial means for DHH individuals to access video and audio content. Improving the accessibility and readability of captions helps create a more inclusive media experience for this group.
- Addressing the speaker indication problem in multi-speaker scenarios with enhanced captioning methods can significantly improve the interactive experience for DHH users.
-
Research Motivation and Related Work:
- The authors identified limitations in existing speaker indication methods that rely primarily on visual cues (e.g., inability to handle frequent speaker switches). Meanwhile, haptic assistive technology has shown potential in audio perception and environmental interaction.
Solution
-
Proposed Method or Solution:
- The authors introduced the "Haptic-Captioning" system, which uses an audio-haptic interface to convert sound directly into tactile vibrations that correspond to sound characteristics (e.g., volume and pitch), enhancing speaker indication in real-time captions.
-
Innovative Aspects:
- Multimodal Feedback: Combines traditional visual captions with haptic feedback to improve caption readability and speaker recognition accuracy.
- No Significant Delay: Converts sound into tactile vibrations instantaneously, avoiding the delay issues of visual caption methods.
- Direct Perception: Simulates the auditory perception of different voice characteristics, enabling DHH users to identify speakers through vibrations.
-
Implementation Steps and Key Technologies:
- Device Design: Developed vibration devices equipped with voice coil actuators, designed to be wearable, handheld, or attachable.
- Experiment Design:
- Preliminary Study: Tested the ability to identify speakers using purely haptic feedback.
- Comparative Study: Compared the system with existing visual captioning methods.
- Contextual Studies: Explored its application across various devices and media types (e.g., movies, live broadcasts, sports events).
- Data Analysis: Collected user feedback, accuracy metrics, and subjective questionnaire responses to quantify the system's effectiveness.
Research Findings
-
Specific Results:
- Preliminary Study Results: Participants achieved a 72.34% accuracy rate in identifying 2-3 speakers using vibrations, demonstrating the potential of haptic feedback for speaker indication.
- Method Comparison Results:
- Haptic-Captioning outperformed traditional real-time captions and some non-real-time visual caption methods in speaker indication accuracy (93.75%).
- User subjective ratings indicated that Haptic-Captioning provided a better overall experience than traditional real-time captions, though it was less preferred compared to some non-real-time visual methods.
- Contextual Study Results:
- Users reported that haptic feedback enhanced caption comprehension, emotional perception, and media engagement.
- Specific scenarios and environmental factors (e.g., multi-person interactions, noisy environments) highlighted the need for more discrete and modular haptic designs.
-
Advantages Compared to Existing Solutions:
- Achieves real-time caption enhancement without significant delays.
- Provides more intuitive multimodal feedback, reducing the cognitive load associated with visual captions.
-
Experimental or Evaluation Results:
- Experiments showed that haptic enhancements helped DHH users perceive captions and speaker switches more smoothly across various scenarios.
- Most users reported that combining haptic and visual information improved the quality of information acquisition.
-
Limitations and Future Directions:
- Limitations:
- Learning Curve: Users need to familiarize themselves with haptic patterns, which may initially cause discomfort or misidentification.
- Adjustability of Vibration Intensity: Limited adjustability may lead to fatigue or sensory overload during prolonged use.
- Public Use Concerns: Potential issues with privacy and noise interference in public environments.
- Future Directions:
- Optimize haptic algorithms to standardize audio-to-haptic conversion.
- Enhance multimodal caption designs by integrating visual and haptic feedback.
- Explore applicability in real-time group conversation scenarios.
- Investigate the impact of external environments (e.g., public transportation) on device performance and propose adaptation strategies.
- Limitations:
The above provides a structured summary and key insights from the paper, facilitating an understanding of the proposed solution and its implications.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- In real-time captions, how can audio-haptic interfaces improve speaker identification for deaf and hard-of-hearing users?Category: Captions, Speech Transcription, and Deaf Communication AssistanceSimilar questionsarrow_forward
- How much does haptic feedback improve speaker identification and information access efficiency in real-time captions?Category: Captions, Speech Transcription, and Deaf Communication AssistanceSimilar questionsarrow_forward
- How do multimodal feedback combinations of haptics and vision help deaf and hard-of-hearing users better understand complex auditory environments?Category: Captions, Speech Transcription, and Deaf Communication AssistanceSimilar questionsarrow_forward
Practical Problems
1- Deaf and hard-of-hearing users struggle to identify speakers through real-time captions, especially in multi-person interactions.Category: Captions, Speech Transcription, and Deaf Communication AssistanceSimilar questionsarrow_forward
- 100%
Tactile Emotions: Multimodal Affective Captioning with Haptics Improves Narrative Engagement for d/Deaf and Hard-of-Hearing Viewers
CHI '25· Vibrotactile Feedback & Skin Stimulation +1
- 60%
Supporting Rhythm Activities of Deaf Children using Music-Sensory-Substitution Systems
CHI '18· Vibrotactile Feedback & Skin Stimulation +1
Based on Jaccard similarity of research subtopics & professions (≥60%)