Sound2Hap: Learning Audio-to-Vibrotactile Haptic Generation from Human Ratings

Best Paper
Vibrotactile Feedback & Skin StimulationAcoustic Sensing & Audio ProcessingAffective Feedback & Emotion Regulation InterfacesUI/UX DesignersAI/ML Researchers & EngineersHCI Researchers

Paper Title

Sound2Hap: Learning Audio-to-Vibrotactile Haptic Generation from Human Ratings

Publication Info

  • Topic area: Audio-to-haptic signal generation using machine learning.
  • Keywords: Audio-to-haptic, vibrotactile feedback, environmental sounds, CNN autoencoder, user ratings, haptic experience, signal processing, generative model, accessibility, interactive tool.

Background and Problem

  • Problem / challenge: Existing audio-to-haptic methods rely on signal-processing rules that are optimized for specific contexts (e.g., music, games) and fail to generalize across diverse environmental sounds. They lack adaptability and perceptual alignment with human preferences.
  • Significance: High-quality audio-to-haptic mapping can enhance realism and accessibility in applications such as virtual reality, gaming, and assistive technologies. However, current methods are limited in their ability to handle diverse and nuanced soundscapes.
  • Motivation and related work: Prior work has focused on signal-processing techniques and machine learning models for audio-to-haptic conversion, but these approaches are either domain-specific or rely on predefined mappings that lack flexibility. This paper seeks to address these limitations with a data-driven approach.

Solution

  • Proposed approach: Sound2Hap, a CNN-based autoencoder model trained on human-rated audio-vibration pairs to generate perceptually meaningful vibrotactile signals from diverse environmental sounds.
  • Novelty:
    1. Creation of the largest user-rated dataset of 4,000 audio-vibration pairs across 50 classes of environmental sounds.
    2. Development of Sound2Hap, a generative model with two training schemes (Top-Pair and Preference-Weighted) for audio-to-vibration translation.
    3. Introduction of an interactive web tool for dataset visualization and vibration generation.
  • Procedure and key techniques:
    1. Implemented four signal-processing algorithms to generate vibrations from 1,000 sound clips in the ESC-50 dataset.
    2. Collected 8,000 user ratings on audio-vibration pairs to create a training dataset.
    3. Trained two Sound2Hap variants: Top-Pair (trained on the best-rated vibration per clip) and Preference-Weighted (trained on a blended target of all rated vibrations).
    4. Conducted user studies to evaluate the model against signal-processing baselines.
    5. Developed a web tool for exploring the dataset and generating vibrations.

Results

  • Concrete findings:
    • Sound2Hap achieved significantly higher audio-vibration match ratings (76.28 for Top-Pair, 73.27 for Preference-Weighted) compared to the baseline (58.28).
    • Both variants outperformed the baseline on the Haptic Experience Index (HXI), with improvements in Harmony, Discord, and General Score.
    • Vibration generation latency was under 1 second for clips up to 20 seconds long.
  • Advantage over baselines:
    • Sound2Hap generalized better to diverse environmental sounds and captured clip-level perceptual nuances, outperforming the best signal-processing algorithm for each sound class.
    • Users preferred Sound2Hap vibrations for their accuracy, rhythm, and intensity.
  • Experiments / evaluation:
    • Study 1: 34 participants rated 4,000 audio-vibration pairs generated by four signal-processing algorithms.
    • Study 2: 15 participants compared Sound2Hap variants against the best-performing signal-processing baseline using 600 new sound clips from ESC-50 and BBC Sound Effects datasets.
    • Metrics included user ratings (0–100 scale) and HXI scores (1–7 scale).
  • Limitations and future work:
    • Dataset limited to single, salient sound sources; future work could address overlapping sounds and broader sound domains.
    • Evaluation focused on finger-based perception; future studies should explore other body locations and actuator types.
    • Sequential playback used for evaluation; concurrent audio-haptic playback should be tested in future research.
    • Demographic imbalance in participant samples; future studies should include more diverse populations.

Summary

Sound2Hap is a CNN-based generative model for converting diverse environmental sounds into perceptually aligned vibrotactile feedback. Trained on a large dataset of human-rated audio-vibration pairs, it outperformed traditional signal-processing methods in user studies, achieving higher ratings for audio-vibration match and haptic experience. The model demonstrated generalizability across sound datasets and maintained low latency. An interactive web tool was developed to support dataset exploration and vibration generation. Future work will address overlapping sounds, multimodal playback, and broader demographic representation. Sound2Hap has potential applications in virtual reality, gaming, and accessibility.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222036/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790649
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
Best Paper
group
Authors
2 authors
sell
Subtopics
Vibrotactile Feedback & Skin Stimulation, Acoustic Sensing & Audio Processing, Affective Feedback & Emotion Regulation Interfaces
work
Professions
UI/UX Designers, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers