Sound2Hap: Learning Audio-to-Vibrotactile Haptic Generation from Human Ratings
Best PaperPaper Title
Sound2Hap: Learning Audio-to-Vibrotactile Haptic Generation from Human Ratings
Publication Info
- Topic area: Audio-to-haptic signal generation using machine learning.
- Keywords: Audio-to-haptic, vibrotactile feedback, environmental sounds, CNN autoencoder, user ratings, haptic experience, signal processing, generative model, accessibility, interactive tool.
Background and Problem
- Problem / challenge: Existing audio-to-haptic methods rely on signal-processing rules that are optimized for specific contexts (e.g., music, games) and fail to generalize across diverse environmental sounds. They lack adaptability and perceptual alignment with human preferences.
- Significance: High-quality audio-to-haptic mapping can enhance realism and accessibility in applications such as virtual reality, gaming, and assistive technologies. However, current methods are limited in their ability to handle diverse and nuanced soundscapes.
- Motivation and related work: Prior work has focused on signal-processing techniques and machine learning models for audio-to-haptic conversion, but these approaches are either domain-specific or rely on predefined mappings that lack flexibility. This paper seeks to address these limitations with a data-driven approach.
Solution
- Proposed approach: Sound2Hap, a CNN-based autoencoder model trained on human-rated audio-vibration pairs to generate perceptually meaningful vibrotactile signals from diverse environmental sounds.
- Novelty:
- Creation of the largest user-rated dataset of 4,000 audio-vibration pairs across 50 classes of environmental sounds.
- Development of Sound2Hap, a generative model with two training schemes (Top-Pair and Preference-Weighted) for audio-to-vibration translation.
- Introduction of an interactive web tool for dataset visualization and vibration generation.
- Procedure and key techniques:
- Implemented four signal-processing algorithms to generate vibrations from 1,000 sound clips in the ESC-50 dataset.
- Collected 8,000 user ratings on audio-vibration pairs to create a training dataset.
- Trained two Sound2Hap variants: Top-Pair (trained on the best-rated vibration per clip) and Preference-Weighted (trained on a blended target of all rated vibrations).
- Conducted user studies to evaluate the model against signal-processing baselines.
- Developed a web tool for exploring the dataset and generating vibrations.
Results
- Concrete findings:
- Sound2Hap achieved significantly higher audio-vibration match ratings (76.28 for Top-Pair, 73.27 for Preference-Weighted) compared to the baseline (58.28).
- Both variants outperformed the baseline on the Haptic Experience Index (HXI), with improvements in Harmony, Discord, and General Score.
- Vibration generation latency was under 1 second for clips up to 20 seconds long.
- Advantage over baselines:
- Sound2Hap generalized better to diverse environmental sounds and captured clip-level perceptual nuances, outperforming the best signal-processing algorithm for each sound class.
- Users preferred Sound2Hap vibrations for their accuracy, rhythm, and intensity.
- Experiments / evaluation:
- Study 1: 34 participants rated 4,000 audio-vibration pairs generated by four signal-processing algorithms.
- Study 2: 15 participants compared Sound2Hap variants against the best-performing signal-processing baseline using 600 new sound clips from ESC-50 and BBC Sound Effects datasets.
- Metrics included user ratings (0–100 scale) and HXI scores (1–7 scale).
- Limitations and future work:
- Dataset limited to single, salient sound sources; future work could address overlapping sounds and broader sound domains.
- Evaluation focused on finger-based perception; future studies should explore other body locations and actuator types.
- Sequential playback used for evaluation; concurrent audio-haptic playback should be tested in future research.
- Demographic imbalance in participant samples; future studies should include more diverse populations.
Summary
Sound2Hap is a CNN-based generative model for converting diverse environmental sounds into perceptually aligned vibrotactile feedback. Trained on a large dataset of human-rated audio-vibration pairs, it outperformed traditional signal-processing methods in user studies, achieving higher ratings for audio-vibration match and haptic experience. The model demonstrated generalizability across sound datasets and maintained low latency. An interactive web tool was developed to support dataset exploration and vibration generation. Future work will address overlapping sounds, multimodal playback, and broader demographic representation. Sound2Hap has potential applications in virtual reality, gaming, and accessibility.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)