Fuzzy Feelings: Arousal’s Interpretive Noise and the Case for Acoustic-Based Haptics

Vibrotactile Feedback & Skin StimulationAudio Accessibility (Captions, Sign Language, Vibration)Affective Feedback & Emotion Regulation InterfacesSpeech-Language Pathologists & AudiologistsAssistive Technology Specialists

Paper Title

Fuzzy Feelings: Arousal’s Interpretive Noise and the Case for Acoustic-Based Haptics

Publication Info

  • Topic area: Accessibility and haptic feedback for Deaf and Hard-of-Hearing individuals.
  • Keywords: haptics, captions, Deaf and Hard-of-Hearing, emotional cues, arousal, acoustic mapping, accessibility, speech-emotion recognition, vibrotactile signals, multimodal design.

Background and Problem

  • Problem / challenge: Captions for Deaf and Hard-of-Hearing (DHH) individuals often fail to convey emotional nuances in speech, particularly arousal. Current approaches to haptic feedback rely on arousal-based mappings, which are ambiguous and inconsistently understood.
  • Significance: Emotional cues in speech are critical for comprehension and engagement, especially for DHH viewers who rely on captions. Addressing this gap can improve accessibility and user experience in media consumption.
  • Motivation and related work: Previous work has explored visual and haptic augmentations to captions, focusing on arousal as a key parameter. However, arousal lacks a stable definition and is often conflated with loudness. Alternative approaches that bypass arousal inference and directly map acoustic features to haptics have shown promise but remain underexplored.

Solution

  • Proposed approach: Acoustic-based haptic mappings that translate pitch, rhythm, and waveform cues into vibrotactile signals, avoiding reliance on arousal inference.
  • Novelty:
    1. Empirical evidence highlighting divergent mental models of arousal among DHH viewers.
    2. Development of novel acoustic-to-haptic mappings for four discrete emotions (happiness, anger, sadness, calmness).
    3. Design guidelines for integrating haptic emotional cues into captioned media.
  • Procedure and key techniques:
    • Part 1: Replicated arousal-driven haptic feedback using speech-emotion recognition (SER) to modulate vibration intensity and typography.
    • Part 2: Tested five acoustic-to-haptic mappings derived from speech features (pitch-normalized, pulse, sawtooth, pitch-exaggerated, unfiltered) across four discrete emotions.
    • Conducted thematic analysis of qualitative responses and quantitative Likert-scale ratings to evaluate interpretability and emotional alignment.

Results

  • Concrete findings:
    • Arousal-based haptics were inconsistently understood, with participants mapping vibrations to loudness rather than emotional arousal.
    • Acoustic-based mappings yielded clearer emotional judgments, with specific patterns aligning with certain emotions (e.g., pulse for high arousal, pitch-normalized for low arousal).
    • Median ratings showed variability across conditions and emotions, with significant contrasts observed in post-hoc tests.
  • Advantage over baselines:
    • Acoustic-based mappings avoided the ambiguity of arousal and provided more interpretable emotional cues.
    • Multi-parameter mappings (e.g., rhythm, pitch, waveform) supported richer affective representation compared to single-parameter intensity mappings.
  • Experiments / evaluation:
    • Participants (n = 14, mean age 43) completed two tasks: evaluating arousal-driven haptics (Part 1) and discrete emotion mappings (Part 2).
    • Stimuli included short video clips and acted utterances, with haptic signals delivered via a wrist-mounted vibrotactile device.
    • Quantitative ratings and qualitative feedback were analyzed to assess interpretability and emotional alignment.
  • Limitations and future work:
    • Small sample size and skewed demographics limit generalizability.
    • Short-form stimuli may not reflect real-world contexts like teleconferencing or public spaces.
    • Functional benefits (e.g., comprehension, recall) were not tested.
    • Future work should compare arousal-based and acoustic-based mappings directly, explore longitudinal effects, and investigate alternative affective constructs.

Summary

This study addresses the challenge of conveying vocal emotion to DHH viewers through haptic feedback. It demonstrates that arousal-based mappings are ambiguous and inconsistently understood, while acoustic-based mappings provide clearer emotional cues by translating pitch, rhythm, and waveform features into vibrotactile signals. Results highlight the importance of multi-parameter designs, cross-modal consistency, and user control in haptic captioning systems. These findings offer actionable guidelines for integrating emotional haptics into short-form media, enhancing accessibility without adding visual load. Future research should explore real-world applications, longitudinal effects, and alternative affective constructs.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222229/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3793421
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Vibrotactile Feedback & Skin Stimulation, Audio Accessibility (Captions, Sign Language, Vibration), Affective Feedback & Emotion Regulation Interfaces
work
Professions
Speech-Language Pathologists & Audiologists, Assistive Technology Specialists
article
Content Status
Full text indexed
hub
Related Papers
0 related papers