Enabling Voice-Accompanying Hand-to-Face Gesture Recognition with Cross-Device Sensing

Honorable Mention
Hand Gesture RecognitionVoice User Interface (VUI) Design

Document Title

Enabling Voice-Accompanying Hand-to-Face Gesture Recognition with Cross-Device Sensing

Document Information

  • Domain: Human-Computer Interaction, Voice Interaction, Multimodal Sensing
  • Keywords: Gesture Recognition, Cross-Device Sensing, Acoustic Sensing, Wearable Devices, Sensor Fusion, Voice Enhancement, Human-Computer Interaction, Hand Movements, Multimodal Interaction

Research Background and Problems

  • Identified Problems or Challenges:

    • Current voice interaction modalities (e.g., wake-up states) face challenges as the modality information in voice is often implicit, and existing natural language processing techniques fail to provide effective support.
    • Users are required to repeat keywords or actively switch target devices, increasing interaction burden.
    • Existing studies primarily focus on single gesture control or fixed voice-gesture schemes, without considering the design of a broader gesture set to enhance voice interaction.
  • Significance:

    • Combining gestures with voice can provide parallel semantic information, helping to expand input channels, simplify voice interface processes, and improve interaction convenience.
    • Hand-to-face gestures are naturally associated with voice, generating significant acoustic features (e.g., blocking sound propagation paths), which can be utilized for sensing and interaction design.
  • Research Motivation and Related Work:

    • Previous studies have demonstrated that parallel body gestures can improve the accuracy and flexibility of voice interaction, but the broader gesture design space and efficient gesture recognition methods based on multi-device sensor fusion remain unexplored.
    • In gesture design, hand-to-face interaction methods have gained attention due to their naturalness and social acceptability, such as the private voice wake-up system PrivateTalk.

Solution

  • Method or Solution:

    • The authors propose an innovative voice-accompanying hand-to-face gesture (VAHF) recognition method based on cross-device sensing. This method combines multiple sensing channels (acoustic, ultrasonic, and inertial sensors) and utilizes commercially available wearable devices such as earbuds, smartwatches, and smart rings for gesture recognition.
    • A user-defined set of 8 VAHF gestures was designed, enabling gesture recognition through cross-device sensor data fusion.
  • Innovations:

    • Developed a novel cross-device sensing technology capable of fusing heterogeneous sensor data from devices like earbuds, smartwatches, and rings.
    • Proposed a recognition model that integrates acoustic sensing, ultrasonic channels, and IMU (Inertial Measurement Unit) data, leveraging deep learning techniques for high-accuracy gesture classification.
    • Expanded the gesture interaction space to support the recognition of up to 8 gestures and their null gestures.
  • Implementation Steps and Key Techniques:

    1. Gesture Design: Conducted user surveys and hierarchical analysis to select 8 VAHF gestures from 15 candidates based on ease of operation, low ambiguity, and high social acceptability.
    2. Data Collection: Collected multi-channel data covering 8 gestures and their accompanying voice using wireless earbuds, smartwatches, and smart rings.
    3. Sensing Scheme: Developed independent acoustic, ultrasonic, and IMU models, combining sensor fusion strategies for gesture classification.
      • The acoustic model extracts spectral features from microphone data.
      • The ultrasonic model estimates hand position using FMCW (Frequency-Modulated Continuous Wave).
      • The IMU model captures hand and finger motion characteristics.
    4. Data Fusion: Explored feature-level and logic-level fusion strategies to enhance model performance.

Research Outcomes

  • Specific Results:

    • Proposed a final VAHF gesture set containing 8 gestures, characterized by good usability, social acceptability, and low fatigue.
    • Achieved 91.5% accuracy for 8-class gesture recognition and 97.3% accuracy for 3-class gesture recognition on a cross-device gesture dataset.
    • Analyzed the differences between each sensor channel and their combinations, contributing to the future design of more efficient sensor fusion schemes.
  • Advantages Compared to Existing Methods:

    • Compared to existing single-device or single-gesture schemes (e.g., PrivateTalk), the proposed method can simultaneously recognize multiple gestures, extending the functionality and scope of voice interaction.
    • Combining multiple sensor types (e.g., acoustic and IMU) ensures stable performance in various environments (e.g., noisy or quiet).
  • Experimental or Evaluation Results:

    • Experiments showed that the feature-level fusion model using the full device combination (earbuds, smartwatch, ring) performed best (91.5% accuracy), while single-earbud configurations significantly reduced accuracy.
    • Simplifying the gesture set to 3 classes improved performance to 97.3%, confirming practical applicability across different environments and device combinations.
  • Limitations and Future Directions:

    • Current experiments were primarily conducted in ideal indoor environments, lacking robustness evaluations in noisy or complex backgrounds.
    • The use of high-frequency ultrasonic waves may raise health and comfort concerns, necessitating further research on safety and exploration of alternative sensing technologies.
    • Future work should optimize hardware deployment (e.g., energy-efficient design), expand the gesture library, and improve models for broader real-world applications.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/96210/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581008
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
Honorable Mention
group
Authors
8 authors
sell
Subtopics
Hand Gesture Recognition, Voice User Interface (VUI) Design
work
Professions
article
Content Status
Full text indexed
hub
Related Papers
5 related papers