OnomaCap: Making Non-speech Sound Captions Accessible and Enjoyable through Onomatopoeic Sound Representation

Voice AccessibilityDeaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)Universal & Inclusive DesignSpeech-Language Pathologists & AudiologistsK-12 TeachersSpecial Education TeachersEarly Childhood Educators

Research Background and Issues

  • What problems or challenges did the authors identify?
    Current non-speech sound captions are primarily described using simple categories (e.g., "[explosion]" or "[laughter]"), neglecting the rich details of sound expression, such as volume changes, temporal variations, or emotional atmosphere. This approach fails to provide a complete auditory experience for Deaf and Hard-of-Hearing (DHH) audiences. Moreover, the process of generating these captions is often time-consuming and inconsistent in quality.

  • Why is this issue important?
    Non-speech sounds (e.g., environmental sounds, sound effects, background music) play a critical role in creating atmosphere and understanding content in videos. Without detailed sound information, DHH audiences may miss important emotional or narrative details.

  • Research Motivation and Related Work
    Inspired by onomatopoeia, which mimics sounds through concise text, the authors believe it can capture the diversity of sound expressions effectively. However, the use of onomatopoeia in non-speech sound captions is limited, and its effectiveness lacks systematic study. This potential motivated the authors to propose a research framework to improve the accessibility of non-speech sounds.

Solution

  • What methods or solutions did the authors propose?
    The authors developed a system called "OnomaCap," which uses an onomatopoeia generation model to transcribe non-speech sounds into onomatopoeic expressions and automatically generate captions. This approach integrates a sound classification module with an onomatopoeia transcription module.

  • What are the innovative aspects of this solution?

    • A specially designed sound transcription model was proposed to directly convert sounds into onomatopoeia.
    • A dataset containing 7,962 sound-onoma pairs was constructed for model training.
    • Detailed user studies were conducted to evaluate the impact of onomatopoeic captions on DHH audiences' video experiences.
  • What are the implementation steps and key technologies used?

    1. Data Collection: Audio samples were extracted from sound datasets and transcribed into onomatopoeia by hearing-abled annotators.
    2. Model Development: Two model architectures were designed (encoder-decoder-based and prefix-tuning models) for onomatopoeia generation.
    3. System Construction: The sound classification model and onomatopoeia transcription model were integrated into a caption generation system.
    4. User Evaluation: Experiments and user interviews were conducted to study the impact of different caption types on DHH audiences' video experiences.

Research Outcomes

  • What specific results were achieved?

    • The onomatopoeia generated by the authors' sound transcription model was subjectively rated as indistinguishable from human-annotated onomatopoeia.
    • DHH participants generally reported that onomatopoeic captions enhanced their understanding of video content and atmosphere, increasing their interest and sense of immersion.
    • Captions combining onomatopoeia with sound categories (category+onoma) were rated as the optimal caption type.
  • What advantages does it have compared to existing solutions?

    1. Compared to traditional category-based captions: Onomatopoeic captions convey more sound details and emotional information.
    2. Compared to graphical sound designs: Onomatopoeic captions maintain high readability and are suitable for various video types.
  • What were the experimental or evaluation results?

    • In user experiments, participants gave subjective ratings of over 3.8 (out of 5) for the onomatopoeia generation model, indicating sufficient naturalness and expressiveness.
    • During the experiments, DHH audiences showed a clear preference for caption types (category, onoma, category+onoma), with category+onoma being the most expressive.
  • Limitations and Future Directions

    • The onomatopoeia generated by the model may be biased due to dataset limitations. For example, this study only used Korean data, and cross-linguistic and cultural generalizability needs further validation.
    • In some cases, using only onomatopoeia may lead to misunderstandings, such as unfamiliar sounds or lack of contextual cues.
    • Future work could explore integrating visual design with onomatopoeic captions to enhance effectiveness and adjusting personalized caption options to meet diverse user needs.

This study significantly advances our understanding of non-speech sound captions and opens new directions for enhancing the audio-visual experiences of DHH audiences.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/189250/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713911
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
2 authors
sell
Subtopics
Voice Accessibility, Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration), Universal & Inclusive Design
work
Professions
Speech-Language Pathologists & Audiologists, K-12 Teachers, Special Education Teachers, Early Childhood Educators
article
Content Status
Full text indexed
hub
Related Papers
2 related papers