ReHEarSSE: Recognizing Hidden-in-the-Ear Silently Spelled Expressions

Electrical Muscle Stimulation (EMS)Augmentative & Alternative Communication (AAC)

Title of the Paper

ReHEarSSE: Recognizing Hidden-in-the-Ear Silently Spelled Expressions

Paper Information

  • Research Domain: Acoustic sensing, artificial intelligence, human-computer interaction (HCI)
  • Keywords: Silent Speech Interaction, Text Input, Autoregressive Model, Earphone Computing, Acoustic Perception

Research Background and Problem Statement

  • Identified Issues or Challenges:
    1. Traditional Silent Speech Interaction (SSI) systems often rely on custom hardware, support only small vocabularies or predefined commands, and struggle to meet broader application scenarios.
    2. SSI systems driven by optical, acoustic, or motion signals face limitations such as device invasiveness, privacy concerns, and noise interference.
    3. Even with ultrasonic sensing in SSI, previous work has failed to expand vocabulary sizes beyond 32-50 words.
  • Significance:
    1. SSI enables private input in quiet environments, avoiding voice interference.
    2. It provides interaction support for individuals unable to speak (e.g., patients with laryngeal separation) or use manual input methods.
  • Motivation and Related Work:
    • The authors propose that active noise cancellation components in earphones can capture ear canal dynamic motion (ECDM) through ultrasonic reflections, potentially supporting large vocabulary SSI.
    • Existing ultrasonic SSI systems, such as EarCommand, rely on custom hardware and have not demonstrated support for large-scale vocabularies or unseen words.

Solution

  • Proposed Method or Solution:
    1. Develop an SSI system named ReHEarSSE, leveraging ultrasonic sensing technology in commercial earphones.
    2. Use autoregressive (AR) features instead of traditional Fourier transform features to capture subtle ear canal changes.
    3. Construct a deep learning model capable of inferring unseen vocabulary using Temporal Convolutional Networks (TCN) and a CTC (Connectionist Temporal Classification) loss function.
  • Innovations:
    1. Supports vocabularies of up to 1000 words and can infer words not present in the training set (i.e., vocabulary generalization capability).
    2. Utilizes optimized Orthogonal Frequency Division Multiplexing (OFDM) signals to enhance ultrasonic signal resolution and information capacity.
    3. Introduces a non-invasive ultrasonic sensing system using earphones, combining privacy and convenience.
  • Implementation Steps and Key Technologies:
    1. Signal Design: Employ optimized OFDM signals instead of traditional FMCW signals to address peak-to-average power ratio (PAPR) issues in conventional signals.
    2. Feature Extraction: Extract AR features and determine the optimal autoregressive order using the Partial Autocorrelation Function (PACF).
    3. Inference Phase: Learn letter embedding features using TCN networks, optimize word prediction with CTC loss, and achieve automatic word matching through ranking scores.

Research Outcomes

  • Specific Results:
    1. Generalization Capability:
      • Achieved an inference accuracy of 89.3% for a vocabulary of 100 unseen words.
      • For larger vocabularies (e.g., 1000 words, with 700 unseen), accuracy reached 73.3%.
    2. Comparison with Existing Solutions:
      • Extended vocabulary size, supported unseen word inference, and implemented using commercial earphones compared to EarCommand.
      • Demonstrated superior performance with OFDM signals over FMCW signals for large vocabulary sizes, showing greater robustness.
    3. Experiments demonstrated system stability under background noise interference, earphone repositioning, and user movement.
  • Experimental or Evaluation Results:
    • Across multiple experimental scenarios (e.g., stationary state, walking, head movement in different directions), the system exhibited robustness, maintaining an average accuracy of 70%-80%.
    • A small amount of personalized user data (e.g., approximately 3-5 minutes of recorded data covering 120 words) significantly improved the model's word inference capability.
  • Limitations and Future Directions:
    1. The initial model requires a certain amount of user training data (e.g., 120 words), posing a slight barrier for new users.
    2. The system has not yet been implemented in real-time; future work should optimize computational efficiency to adapt to earphone edge devices.
    3. Enhancing robustness against background motion (e.g., walking, head displacement) remains a challenge.
    4. Exploring binaural collaboration (dual-ear audio input) may further improve system performance and warrants future investigation.

Conclusion

ReHEarSSE leverages ultrasonic sensing in commercial earphones to achieve non-invasive, large vocabulary support for silent speech input. It represents a significant advancement in the field of silent speech interaction, combining privacy, high generalization capability, and usability. The system shows promise for diverse commercial and medical application scenarios.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147342/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642095
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Electrical Muscle Stimulation (EMS), Augmentative & Alternative Communication (AAC)
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
0 related papers