SPECTRA: Personalizable Sound Recognition for Deaf and Hard of Hearing Users through Interactive Machine Learning

Electrical Muscle Stimulation (EMS)Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)Speech-Language Pathologists & AudiologistsAssistive Technology Specialists

Research Background and Problems

  • What problems or challenges did the authors identify?
    The authors pointed out that many Deaf and Hard of Hearing (DHH) users desire more accurate sound recognition tools capable of identifying a wide range of sound categories to meet needs related to personal safety, social participation, and daily tasks. However, these sound recognition technologies lack personalized adaptation to users' local sound environments, leading to insufficient accuracy. Additionally, a major challenge lies in how to support and guide DHH users in collecting effective audio data and selecting appropriate datasets for training machine learning models, given that these users cannot fully access sound content.

  • Why is this problem important?
    Sound plays a critical role in many real-world scenarios, such as identifying alarms, interacting in social environments, and monitoring household appliances. Therefore, sound recognition technology has the potential to be life-changing for DHH users. However, current solutions fail to meet the diverse needs of DHH users due to the lack of integration between personalized requirements and the complexity of sound recognition.

  • Research Motivation and Related Work
    Previous work has begun exploring the concept of customized sound recognition but still lacks systems capable of providing a complete end-to-end solution, especially for non-auditory users. Additionally, existing commercial devices offer limited sound customization options, failing to provide the transparency and sense of control that users need.

Solution

  • What methods or solutions did the authors propose?
    The authors proposed an interactive machine learning (IML) pipeline called “SPECTRA,” specifically designed for DHH users to personalize sound recognition models. SPECTRA supports users through a three-stage process: recording data, editing training datasets, and generating models, as well as real-time model performance evaluation. Additionally, SPECTRA incorporates several key interface features, including audio visualization through waveforms and spectrograms, interactive clustering visualization of datasets, and the ability for users to add textual annotations.

  • What are the innovative aspects of this solution?

    • It provides the first complete end-to-end personalized sound recognition system for DHH users, covering all stages from recording to training and testing.
    • It introduces an interactive dataset clustering visualization tool, enabling users to explore and optimize audio datasets.
    • It enhances users' transparency and understanding of training datasets, allowing iterative filtering and updating of data.
    • It offers real-time predictions and confidence scores, enabling users to evaluate model performance.
  • What are the implementation steps and key technologies used?

    1. Recording and Dataset Creation: Users monitor environmental sounds in real-time through waveform and spectrogram visualizations and record audio. Each recording is segmented into 1-second clips, and Mel spectrograms are automatically generated for subsequent model use.
    2. Data Filtering and Model Training: Users select high-quality audio clips for training through interactive clustering graphs and waveforms, while excluding noisy or erroneous samples.
    3. Real-Time Testing and Evaluation: Users test the model in real-world environments and observe real-time outputs, including predicted confidence scores and a history of past recognitions.
      Key technologies used include the pre-trained Speech Commands API in TensorFlow.js, UMAP for dataset dimensionality reduction, and multi-dimensional visualization tools (e.g., ScatterGL).

Research Outcomes

  • What specific outcomes were achieved?

    • Through evaluation, all participants successfully trained and tested personalized sound recognition models during the experiment. Users expressed strong support for the clustering visualization tool, noting that it significantly improved their understanding and control over model data.
    • The data annotation feature helped participants capture and distinguish complex variations in sounds, improving data quality.
    • Participants showed a positive attitude toward personalized sound recognition technology and proposed specific future application scenarios.
  • What advantages does it have compared to existing solutions?

    • The SPECTRA system addresses the needs of non-auditory users by integrating visual tools, overcoming the limitations of traditional sound recognition tools, and enabling users to actively participate in data collection and model optimization.
    • It provides a complete interactive training process, offering greater flexibility and applicability compared to existing commercial solutions (e.g., Android and iOS platforms).
  • What were the experimental or evaluation results?

    • Most of the twelve participants reported that SPECTRA effectively guided them through the training and simulation processes for sound recognition, including the use of waveforms and clustering tools.
    • Participants successfully created models capable of recognizing up to six sound categories in home environments and provided detailed quantitative and qualitative data to support the findings.
  • Limitations and Future Directions

    • Limitations include the inability of users to perform complex multiple model retraining within a fixed time frame, and the results have not been tested over long periods in real-world environments.
    • SPECTRA relies on relatively small Mel spectrogram inputs and limited data processing capabilities. Future work could integrate it with more powerful pre-trained frameworks (e.g., ProtoSound) to optimize execution efficiency.
    • The system needs further exploration to improve performance in complex acoustic environments and should enhance lightweight mobile adaptation for daily use.

Through this study, the authors successfully validated the importance of interactive machine learning in personalized sound recognition systems for non-auditory users, providing a strong foundation for further optimization and expansion in this field.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/189442/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713294
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Electrical Muscle Stimulation (EMS), Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille), Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration)
work
Professions
Speech-Language Pathologists & Audiologists, Assistive Technology Specialists
article
Content Status
Full text indexed
hub
Related Papers
5 related papers