MAF: Exploring Mobile Acoustic Field for Hand-to-Face Gesture Interactions

Haptic WearablesHand Gesture Recognition

Document Title

MAF: Exploring Mobile Acoustic Field for Hand-to-Face Gesture Interactions

Document Information

  • Research Area: Human-Computer Interaction, Gesture Recognition, Acoustic Sensing
  • Keywords: Wearable Computing, Gesture Detection, Acoustic Sensing, Surface Acoustic Waves, Bone Conduction Headphones, Deep Learning, User Experience

Research Background and Problem Statement

  • Problems and Challenges:
    • Current gesture detection systems primarily rely on inertial measurement units (IMUs), capacitive sensors, or cameras, which have limited capability for non-contact gesture detection.
    • Camera-based solutions are susceptible to lighting conditions, privacy concerns, and occlusion issues; radar technology is costly and power-intensive; other methods like vibration sensors or acoustic signals require specialized hardware support.
  • Research Importance:
    • Developing a technology capable of detecting facial gestures using conventional and readily available bone conduction headphones could significantly advance gesture-based human-computer interaction in everyday applications and facilitate integration with extended reality (XR) technologies.
  • Research Motivation:
    • The hypothesis is that acoustic-based gesture detection can be achieved without specialized hardware, enabling both contact and non-contact gesture detection.
    • Bone conduction headphones can generate surface acoustic waves (SAWs) and leaky surface acoustic waves (LSAWs), which together form a Mobile Acoustic Field (MAF), presenting new opportunities for gesture detection.

Solution

  • Methods and Techniques:
    • Propose a "Mobile Acoustic Field" (MAF) based on bone conduction headphones, generating surface acoustic waves and leaky acoustic waves. By detecting waveform interference caused by gestures, the system identifies gesture actions.
    • Design a signal processing pipeline including filtering, signal enhancement, segmentation based on KL divergence, and gesture classification.
    • Use a Convolutional Recurrent Neural Network (CRNN) for acoustic signal feature learning and classification.
  • Implementation Steps:
    1. Signal Generation: Emit probing signals in the ultrasonic frequency band using bone conduction headphones to generate stable acoustic waves.
    2. Signal Preprocessing: Apply filters for noise reduction and use Wiener filtering to enhance signal quality.
    3. Signal Segmentation: Detect gesture intervals and segment signals based on KL divergence of the waveform.
    4. Classification and Recognition: Employ a deep learning model combining CNN and LSTM to classify gesture types.

Research Outcomes

  • Specific Results:
    • Successfully designed the signal processing pipeline for the MAF system and identified 10 types of gestures, covering contact gestures like "single facial press" and non-contact gestures like "palm approaching the face."
    • Achieved an average gesture recognition accuracy of 92% in experiments involving 22 participants.
  • Experimental and Evaluation Results:
    • Successfully detected gestures under varying conditions, including different volumes, distances, and activity states (stationary, walking, running).
    • Compared the performance of various machine learning models (SVM, kNN, decision tree, CNN, etc.) and found that CRNN significantly outperformed traditional models in accuracy.
    • Demonstrated negligible performance variation between morning (moist skin) and evening (oily skin) conditions.
    • Showed minimal impact of noise environments, headphone position adjustments, and music playback on recognition accuracy.
  • Advantages and Innovations:
    • Utilizes portable headphones without requiring specialized sensors, offering low cost and easy deployment.
    • Supports gesture detection in complex everyday environments (e.g., walking, speaking).
    • Enables simultaneous gesture detection and audio playback (e.g., speech or music) without interference.
  • Limitations and Future Directions:
    • Currently supports only bone conduction headphones; further research is needed to enable compatibility with more general devices.
    • Users need to attach the microphone to the head, which is inconvenient for daily use. Future exploration could focus on directly utilizing built-in microphones in headphones.
    • The current model performs poorly in recognizing complex micro-gestures (e.g., two-finger dragging or ear pinching). Increasing model capacity could address this but requires balancing processing latency.
    • Future research could explore more advanced algorithms to address motion interference (e.g., recognition impact during walking or running).

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147716/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642437
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Haptic Wearables, Hand Gesture Recognition
work
Professions
article
Content Status
Full text indexed
hub
Related Papers
10 related papers