ExpresSense: Exploring a Standalone Smartphone to Sense Engagement of Users from Facial Expressions Using Acoustic Sensing

Human Pose & Activity RecognitionBiosensors & Physiological Monitoring

Title of the Paper

ExpresSense: Exploring a Standalone Smartphone to Sense Engagement of Users from Facial Expressions Using Acoustic Sensing

Paper Information

  • Research Area: Human-Computer Interaction, Acoustic-Based Facial Expression Detection
  • Keywords: Acoustic sensing, smartphone, expressions, engagement, assistive systems, machine learning, non-visual detection, real-time processing

Research Background and Problem Statement

  • Problem Description:

    • Traditional image- and video-based facial expression detection methods face limitations such as strong dependency on lighting conditions, detection failures due to occlusion, privacy concerns, and high computational and energy demands.
    • Meeting the requirements for real-time facial expression detection on resource-constrained devices like smartphones poses significant challenges.
  • Importance:

    • Detecting facial expressions associated with user emotions and engagement can enable applications in mental health monitoring and emotion management, such as early detection of depression symptoms or enhancing user experience.
    • Non-visual facial expression detection in privacy-sensitive, energy-efficient, and interference-prone scenarios has broad practical potential.
  • Research Motivation and Related Work:

    • Current approaches (e.g., SonicFace, EarIO) for acoustic facial expression detection rely on external hardware (e.g., microphone arrays or earphones) or special setups.
    • This study aims to develop a lightweight solution that can be implemented using consumer-grade smartphones alone, eliminating dependency on external hardware and addressing current limitations in acoustic detection.

Solution

  • Method/Solution:

    • Propose an acoustic sensing system—ExpresSense—that utilizes the built-in speaker and microphone of smartphones to generate 16kHz-19kHz near-ultrasound signals and capture phase and amplitude changes in facial reflection signals to detect basic expressions.
  • Innovations:

    1. Utilizes existing hardware in commercial smartphones without requiring external devices.
    2. Introduces lightweight real-time signal processing and machine learning models to balance accuracy and performance.
    3. Operates under various lighting conditions without relying on privacy-sensitive camera data.
    4. Overcomes detection challenges in occlusion scenarios (e.g., wearing glasses or masks).
  • Implementation Steps and Key Techniques:

    1. Signal Generation and Reception: Generate linear frequency-modulated continuous wave (FMCW) signals between 16-19kHz and receive reflected signals via the smartphone.
    2. Signal Processing:
      • Use high-pass filters to eliminate low-frequency noise interference.
      • Extract frequency-domain features using Fourier transform and select key frequency components relevant to facial regions.
      • Remove static reflection interference (e.g., from walls and tables).
    3. Expression Classification:
      • Extract amplitude and phase features from reflected signals.
      • Perform expression classification using three classifiers (Logistic Regression, Decision Tree, Random Forest) and generate final results through majority voting.
    4. System Optimization:
      • Design an Android application for experiments and user testing.
      • Conduct sensitivity analysis under varying experimental conditions (e.g., environmental noise, movement, device angle changes).

Research Outcomes

  • Specific Results:

    • Developed a facial expression detection tool that operates without a camera and is entirely smartphone-based.
    • Achieved an average classification accuracy of approximately 75%, offering significant simplicity compared to other acoustic expression detection technologies.
  • Comparison with Existing Solutions:

    • Significantly reduced hardware requirements: no need for microphone arrays, external earphones, or additional devices.
    • Demonstrated robust performance in specific scenarios (e.g., low-light conditions, occlusion).
  • Experimental or Evaluation Results:

    • Experimental Results:
      • Achieved an average accuracy of approximately 75% with user-dependent models, evaluated through multidimensional tests (e.g., angle deviations, environmental noise, motion effects).
      • Performed well in natural usage scenarios, such as personalized user engagement scoring for streaming content evaluation tasks, achieving an average F1 score of 0.84.
    • Sensitivity Analysis:
      • Evaluated model robustness under varying conditions such as distance, phone tilt angle, noise interference, and handheld methods.
      • While environmental motion and finger dynamics had limited impact on overall performance, significant motion or device occlusion did reduce detection accuracy.
  • Limitations and Future Directions:

    1. Sound Applicability: The current 16-19kHz signals may be slightly audible in certain cases, requiring further optimization of signal range or improved device frequency support (e.g., extension to iOS devices).
    2. Dynamic Scenarios: Calibration and optimization are needed for scenarios involving significant user movement (e.g., walking) or frequent device motion.
    3. Misclassification and Generalization: Accuracy declines for new users; future work should focus on building larger datasets for training to improve generalizability.
    4. Expression Accuracy Impact: Classification rates for specific expressions (e.g., sadness, surprise) are relatively low; further sensitivity analysis of AU (Action Unit) distribution is needed.

Conclusion

ExpresSense demonstrates the potential of acoustic-based facial expression detection without cameras, suitable for a range of real-time user emotion detection tasks. It provides theoretical and practical insights for privacy-friendly and lightweight next-generation smartphone applications.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/96564/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581235
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Human Pose & Activity Recognition, Biosensors & Physiological Monitoring
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
10 related papers