SpeechCompass: Enhancing Mobile Captioning with Diarization and Directional Guidance via Multi-Microphone Localization
Best PaperAuthors
Voice AccessibilityBiosensors & Physiological MonitoringDisability Service ProvidersAssistive Technology Specialists
Research Background and Problem Statement
-
Identified Problems/Challenges:
- Current mobile real-time speech-to-text (ASR) applications struggle to effectively distinguish between speakers in multi-speaker scenarios, leading to chaotic transcription.
- All utterances are often concatenated into a single text stream, failing to reflect speaker direction and identity, which hinders users from understanding speaker transitions in conversations.
- Background noise and nearby irrelevant conversations are also transcribed, causing additional confusion and privacy concerns.
-
Why the Problem is Important:
- Speech-to-text technology has significant applications in hearing assistance, language translation, and meeting transcription. However, its inability to distinguish speakers or provide directional information limits its potential.
- In multi-speaker conversations, users need clear identification of speakers or their directions to better comprehend the dialogue.
-
Research Motivation and Related Work:
- The authors' survey shows that the majority of users (60%) frequently encounter challenges with speech-to-text technology in separating multiple speakers and handling background noise.
- Traditional single-microphone or simple embedded microphone array methods fail to meet the demands of real-time separation and directional display.
- While microphone array processing has been applied to speaker separation in conference rooms, it has not been extended to mobile devices or everyday environments.
Proposed Solution
-
Proposed Method or Solution:
- Develop a system called "SpeechCompass" that utilizes a four-microphone array and a low-latency directional localization algorithm to achieve real-time speech separation and directional display.
- This includes developing a hardware prototype in the form of a smartphone case (integrated with a four-microphone array) and a mobile application with visualization tools to help users identify speaker directions.
-
Innovative Aspects of the Solution:
- Introduced a real-time audio localization algorithm that operates on low-power microcontrollers, enabling 360° speaker direction localization.
- Integrated the microphone array with mobile devices to support a portable real-time ASR solution, reducing computational costs for processing directional information while ensuring privacy and low latency.
- Provided various visualization options, such as text color indicating speaker direction, on-screen arrows, or a mini-map to show the sound source location.
-
Implementation Steps and Key Technologies:
- Hardware: Designed a smartphone case integrated with four microphones and used an STM32L55 low-power microcontroller for data collection and audio processing.
- Algorithm: Employed an improved Generalized Cross-Correlation with Phase Transform (GCC-PHAT) algorithm to efficiently calculate Time Difference of Arrival (TDOA) between microphones and determine direction.
- User Interface: Developed an Android application to enable real-time subtitles and directional visualization, offering different presentation modes (e.g., arrows, text color, mini-map indicating speaker direction).
- Experiments and Evaluation: Combined user survey data to evaluate the system's performance and user feedback in multi-speaker, noisy environments.
Research Outcomes
-
Specific Achievements:
- Developed an embedded four-microphone prototype hardware capable of accurately localizing speaker direction in real-time using local algorithms, seamlessly integrated with mobile devices for an enhanced ASR experience.
- Proposed an interactive UI design that distinguishes speaker identity and direction in multi-speaker conversations.
-
Advantages Over Existing Solutions:
- Improved transcription readability: speaker separation achieved via directional indication.
- Compared to solutions relying solely on machine learning, the microphone array approach offers lower computational costs and higher privacy.
- Supports simplified configurations, such as 180° direction detection using a dual-microphone setup.
-
Experimental or Evaluation Results:
- The microphone array achieved a directional measurement error of 11.1° to 22.1°, approaching human-level directional recognition, particularly at normal conversation volumes (≥60 dB).
- Social surveys revealed that enhanced speaker differentiation in group conversations increased user satisfaction and willingness to recommend the subtitle application.
- In experiments, the UI combining arrow-based directional display with text color was the most favored by users.
-
Limitations and Future Directions:
- Limitations:
- Elevation angles (speakers positioned higher or lower) can affect directional measurement accuracy.
- Background noise interferes with the localization algorithm, potentially impacting accuracy.
- Future Directions:
- Expand the microphone array topology design, such as integration into wearable devices like smartwatches or glasses.
- Leverage machine learning to enhance noise robustness and improve directional measurement.
- Conduct large-scale, long-term user testing to explore the system's adaptability to various user scenarios in real-world environments.
- Limitations:
The above analysis presents a comprehensive overview of SpeechCompass, offering an innovative solution to the multi-speaker group subtitle problem while paving the way for future advancements in mobile speech applications.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How can speaker diarization be achieved for real-time mobile speech-to-text (ASR) in multi-speaker scenes?Category: Noise-Robust Speech Capture DevicesSimilar questionsarrow_forward
- Can microphone-array technology and directional visualization improve ASR usability and accuracy?Category: Noise-Robust Speech Capture DevicesSimilar questionsarrow_forward
- How effective are four-microphone arrays and low-latency direction-finding algorithms in practice?Category: Noise-Robust Speech Capture DevicesSimilar questionsarrow_forward
lightbulb
Practical Problems
1- In multi-speaker conversations, users struggle to distinguish speakers and transcripts become confused.Category: Noise-Robust Speech Capture DevicesSimilar questionsarrow_forward
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713631
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
Best Paper
group
Authors
6 authors
sell
Subtopics
Voice Accessibility, Biosensors & Physiological Monitoring
work
Professions
Disability Service Providers, Assistive Technology Specialists
article
Content Status
Full text indexed
hub
Related Papers
1 related papers