SpeechCompass: Enhancing Mobile Captioning with Diarization and Directional Guidance via Multi-Microphone Localization

Best Paper
Voice AccessibilityBiosensors & Physiological MonitoringDisability Service ProvidersAssistive Technology Specialists

Research Background and Problem Statement

  • Identified Problems/Challenges:

    • Current mobile real-time speech-to-text (ASR) applications struggle to effectively distinguish between speakers in multi-speaker scenarios, leading to chaotic transcription.
    • All utterances are often concatenated into a single text stream, failing to reflect speaker direction and identity, which hinders users from understanding speaker transitions in conversations.
    • Background noise and nearby irrelevant conversations are also transcribed, causing additional confusion and privacy concerns.
  • Why the Problem is Important:

    • Speech-to-text technology has significant applications in hearing assistance, language translation, and meeting transcription. However, its inability to distinguish speakers or provide directional information limits its potential.
    • In multi-speaker conversations, users need clear identification of speakers or their directions to better comprehend the dialogue.
  • Research Motivation and Related Work:

    • The authors' survey shows that the majority of users (60%) frequently encounter challenges with speech-to-text technology in separating multiple speakers and handling background noise.
    • Traditional single-microphone or simple embedded microphone array methods fail to meet the demands of real-time separation and directional display.
    • While microphone array processing has been applied to speaker separation in conference rooms, it has not been extended to mobile devices or everyday environments.

Proposed Solution

  • Proposed Method or Solution:

    • Develop a system called "SpeechCompass" that utilizes a four-microphone array and a low-latency directional localization algorithm to achieve real-time speech separation and directional display.
    • This includes developing a hardware prototype in the form of a smartphone case (integrated with a four-microphone array) and a mobile application with visualization tools to help users identify speaker directions.
  • Innovative Aspects of the Solution:

    • Introduced a real-time audio localization algorithm that operates on low-power microcontrollers, enabling 360° speaker direction localization.
    • Integrated the microphone array with mobile devices to support a portable real-time ASR solution, reducing computational costs for processing directional information while ensuring privacy and low latency.
    • Provided various visualization options, such as text color indicating speaker direction, on-screen arrows, or a mini-map to show the sound source location.
  • Implementation Steps and Key Technologies:

    1. Hardware: Designed a smartphone case integrated with four microphones and used an STM32L55 low-power microcontroller for data collection and audio processing.
    2. Algorithm: Employed an improved Generalized Cross-Correlation with Phase Transform (GCC-PHAT) algorithm to efficiently calculate Time Difference of Arrival (TDOA) between microphones and determine direction.
    3. User Interface: Developed an Android application to enable real-time subtitles and directional visualization, offering different presentation modes (e.g., arrows, text color, mini-map indicating speaker direction).
    4. Experiments and Evaluation: Combined user survey data to evaluate the system's performance and user feedback in multi-speaker, noisy environments.

Research Outcomes

  • Specific Achievements:

    • Developed an embedded four-microphone prototype hardware capable of accurately localizing speaker direction in real-time using local algorithms, seamlessly integrated with mobile devices for an enhanced ASR experience.
    • Proposed an interactive UI design that distinguishes speaker identity and direction in multi-speaker conversations.
  • Advantages Over Existing Solutions:

    • Improved transcription readability: speaker separation achieved via directional indication.
    • Compared to solutions relying solely on machine learning, the microphone array approach offers lower computational costs and higher privacy.
    • Supports simplified configurations, such as 180° direction detection using a dual-microphone setup.
  • Experimental or Evaluation Results:

    • The microphone array achieved a directional measurement error of 11.1° to 22.1°, approaching human-level directional recognition, particularly at normal conversation volumes (≥60 dB).
    • Social surveys revealed that enhanced speaker differentiation in group conversations increased user satisfaction and willingness to recommend the subtitle application.
    • In experiments, the UI combining arrow-based directional display with text color was the most favored by users.
  • Limitations and Future Directions:

    • Limitations:
      • Elevation angles (speakers positioned higher or lower) can affect directional measurement accuracy.
      • Background noise interferes with the localization algorithm, potentially impacting accuracy.
    • Future Directions:
      • Expand the microphone array topology design, such as integration into wearable devices like smartwatches or glasses.
      • Leverage machine learning to enhance noise robustness and improve directional measurement.
      • Conduct large-scale, long-term user testing to explore the system's adaptability to various user scenarios in real-world environments.

The above analysis presents a comprehensive overview of SpeechCompass, offering an innovative solution to the multi-speaker group subtitle problem while paving the way for future advancements in mobile speech applications.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/189502/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713631
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
Best Paper
group
Authors
6 authors
sell
Subtopics
Voice Accessibility, Biosensors & Physiological Monitoring
work
Professions
Disability Service Providers, Assistive Technology Specialists
article
Content Status
Full text indexed
hub
Related Papers
1 related papers