ProtoSound: A Personalized and Scalable Sound Recognition System for Deaf and Hard-of-Hearing Users
Authors
Document Title
ProtoSound: A Personalized and Scalable Sound Recognition System for Deaf and Hard-of-Hearing Users
Document Information
- Domain: Human-Computer Interaction and Machine Learning
- Keywords: Customizable Sound Recognition, Few-Shot Learning, Technology for the Deaf, Environmental Sound Recognition, Sound Perception, Mobile Applications
Research Background and Problem Statement
-
What problems or challenges did the authors identify?
- Current sound recognition systems rely on pre-trained generic models, which fail to meet the diverse needs of hearing-impaired users.
- Existing systems lack support for user-customized sound categories, such as household-specific sounds (e.g., a child’s voice or a pet’s bark).
- Generic models cannot adapt to the diversity of audio environments in daily life, such as transitions between indoor and outdoor settings.
-
Why is this problem important?
- Sound recognition systems can provide critical alerts about the environment, activities, and emergencies for hearing-impaired users, improving their quality of life and safety.
- Personalization and environmental adaptability can significantly enhance system accuracy and user satisfaction.
-
Motivation and related work:
- Based on previous research and a survey of 472 hearing-impaired participants, the authors identified a strong demand for personal sound recognition technology.
- The authors conducted an in-depth analysis of the limitations of existing sound recognition tools and their shortcomings in real-world scenarios.
- Inspired by prior research on few-shot learning, the authors proposed a more user-friendly system design and implementation.
Solution
-
What methods or solutions did the authors propose?
- ProtoSound System: A few-shot learning model that customizes sound categories based on a small number of user-provided samples.
- Real-Time Personalized Model: Enables real-time model customization on mobile devices using a small amount of audio recordings.
- Context Generalization Techniques: Introduces data augmentation methods to adapt the model to background noise in different environments.
- Open-Set Recognition: Allows the system to identify unknown sound categories.
-
What are the innovative aspects of this solution?
- Proposed the first scalable sound recognition system with low user involvement costs.
- The system design integrates the technical features of few-shot learning while addressing personalization and real-time operation needs.
- Supports the recognition of hard-to-record target sound categories by enhancing recognition capabilities with a built-in online resource library.
-
What are the implementation steps? What key technologies were used?
- Data Collection: Users record a small number of samples for model training, requiring approximately 5 samples per category.
- Model Training: Features are extracted using the MobileNetV2 architecture, and category prototypes are generated through a prototypical network.
- Prediction Phase: Compares test audio with category prototypes and uses a nearest-neighbor classifier to predict sound categories.
- Enhanced Features: Includes context generalization (via background noise data augmentation), open-set classification (recognizing unknown categories), and real-time end-to-end application deployment.
Research Outcomes
-
What specific results were achieved?
- Developed the ProtoSound system, implemented through open-source Python code and an Android application.
- Conducted experiments on two real-world datasets, demonstrating that the system outperforms other methods in accuracy.
- Field studies showed that ProtoSound can handle diverse environments and personalized sound categories, with an average user recording time of 10 minutes and a prediction accuracy of 87.4%.
-
What advantages does it have compared to existing solutions?
- Enables real-time deployment of few-shot learning, allowing users to train models on mobile devices without requiring extensive computational resources.
- Supports personalization, addressing the limitations of generic models in handling user-specific sounds.
- Achieves higher accuracy (+9.7%) compared to current few-shot learning and traditional supervised learning methods.
-
What were the experimental or evaluation results?
- Experiment 1: Achieved 90.4% accuracy on sound data from hearing-impaired participants, significantly outperforming supervised learning baseline models.
- Experiment 2: Achieved accuracy close to human labeling (average 91.3%) on detected sounds in real-world scenarios.
- Experiment 3: Field studies demonstrated that ProtoSound operates effectively in various locations (e.g., homes, restaurants, streets), with positive user feedback.
-
Limitations and future directions:
- Limited Categories: Current experiments focus on a 5-way setup (5 categories); future work could expand to more categories.
- User Interface Improvements: Further optimization is needed for the interface used for sound recording and labeling, especially for hearing-impaired users.
- Long-Term Deployment Evaluation: Current research focuses on short-term evaluations; future studies should assess performance over extended periods and across multiple scenarios.
- Socio-Cultural Impact: Consider the diverse needs and cultural preferences of hearing-impaired user groups, including their acceptance of sound perception technology.
Open-Source Resources and Contributions
- Python and Android open-source implementation: https://github.com/makeabilitylab/ProtoSound
- Integration of open sound libraries and user interaction design.
- Provided a dataset and evaluation methods to support multi-scenario research needs.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can a personalized and extensible sound recognition system be designed to meet deaf and hard-of-hearing users' needs?Category: Deaf and Hard-of-Hearing ASL and Caption SupportSimilar questionsarrow_forward
- How can few-shot learning be used to enable personalized sound category customization from user-provided samples?Category: Deaf and Hard-of-Hearing ASL and Caption SupportSimilar questionsarrow_forward
- How can sound recognition systems adapt to background noise in different environments and recognize unknown categories?Category: Deaf and Hard-of-Hearing ASL and Caption SupportSimilar questionsarrow_forward
Practical Problems
1- Deaf and hard-of-hearing users struggle to receive safety- and convenience-related environmental sound alerts.Category: Deaf and Hard-of-Hearing ASL and Caption SupportSimilar questionsarrow_forward
- 80%
Write-it-Yourself with the Aid of Smartwatches: A Wizard-of-Oz Experiment with Blind People
IUI '18· Deaf & Hard-of-Hearing Support (Captions, Sign Language, Vibration) +1
- 67%
Methods for Evaluation of Imperfect Captioning Tools by Deaf or Hard-of-Hearing Users at Different Reading Literacy Levels
CHI '18· Voice Accessibility +1
- 67%
TacNote: Tactile and Audio Note-Taking for Non-Visual Access
UIST '23· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille) +2
- 60%
Stroke-Gesture Input for People with Motor Impairments: Empirical Results & Research Roadmap
CHI '19· Motor Impairment Assistive Input Technologies
- 60%
PersonalTouch: Improving Touchscreen Usability by Personalizing Accessibility Settings based on Individual User's Touchscreen Interaction
CHI '19· Motor Impairment Assistive Input Technologies
- 60%
A Performance Evaluation of Nomon: A Flexible Interface for Noisy Single-Switch Users
CHI '22· Motor Impairment Assistive Input Technologies
Based on Jaccard similarity of research subtopics & professions (≥60%)