ExpresSense: Exploring a Standalone Smartphone to Sense Engagement of Users from Facial Expressions Using Acoustic Sensing
Authors
Title of the Paper
ExpresSense: Exploring a Standalone Smartphone to Sense Engagement of Users from Facial Expressions Using Acoustic Sensing
Paper Information
- Research Area: Human-Computer Interaction, Acoustic-Based Facial Expression Detection
- Keywords: Acoustic sensing, smartphone, expressions, engagement, assistive systems, machine learning, non-visual detection, real-time processing
Research Background and Problem Statement
-
Problem Description:
- Traditional image- and video-based facial expression detection methods face limitations such as strong dependency on lighting conditions, detection failures due to occlusion, privacy concerns, and high computational and energy demands.
- Meeting the requirements for real-time facial expression detection on resource-constrained devices like smartphones poses significant challenges.
-
Importance:
- Detecting facial expressions associated with user emotions and engagement can enable applications in mental health monitoring and emotion management, such as early detection of depression symptoms or enhancing user experience.
- Non-visual facial expression detection in privacy-sensitive, energy-efficient, and interference-prone scenarios has broad practical potential.
-
Research Motivation and Related Work:
- Current approaches (e.g., SonicFace, EarIO) for acoustic facial expression detection rely on external hardware (e.g., microphone arrays or earphones) or special setups.
- This study aims to develop a lightweight solution that can be implemented using consumer-grade smartphones alone, eliminating dependency on external hardware and addressing current limitations in acoustic detection.
Solution
-
Method/Solution:
- Propose an acoustic sensing system—ExpresSense—that utilizes the built-in speaker and microphone of smartphones to generate 16kHz-19kHz near-ultrasound signals and capture phase and amplitude changes in facial reflection signals to detect basic expressions.
-
Innovations:
- Utilizes existing hardware in commercial smartphones without requiring external devices.
- Introduces lightweight real-time signal processing and machine learning models to balance accuracy and performance.
- Operates under various lighting conditions without relying on privacy-sensitive camera data.
- Overcomes detection challenges in occlusion scenarios (e.g., wearing glasses or masks).
-
Implementation Steps and Key Techniques:
- Signal Generation and Reception: Generate linear frequency-modulated continuous wave (FMCW) signals between 16-19kHz and receive reflected signals via the smartphone.
- Signal Processing:
- Use high-pass filters to eliminate low-frequency noise interference.
- Extract frequency-domain features using Fourier transform and select key frequency components relevant to facial regions.
- Remove static reflection interference (e.g., from walls and tables).
- Expression Classification:
- Extract amplitude and phase features from reflected signals.
- Perform expression classification using three classifiers (Logistic Regression, Decision Tree, Random Forest) and generate final results through majority voting.
- System Optimization:
- Design an Android application for experiments and user testing.
- Conduct sensitivity analysis under varying experimental conditions (e.g., environmental noise, movement, device angle changes).
Research Outcomes
-
Specific Results:
- Developed a facial expression detection tool that operates without a camera and is entirely smartphone-based.
- Achieved an average classification accuracy of approximately 75%, offering significant simplicity compared to other acoustic expression detection technologies.
-
Comparison with Existing Solutions:
- Significantly reduced hardware requirements: no need for microphone arrays, external earphones, or additional devices.
- Demonstrated robust performance in specific scenarios (e.g., low-light conditions, occlusion).
-
Experimental or Evaluation Results:
- Experimental Results:
- Achieved an average accuracy of approximately 75% with user-dependent models, evaluated through multidimensional tests (e.g., angle deviations, environmental noise, motion effects).
- Performed well in natural usage scenarios, such as personalized user engagement scoring for streaming content evaluation tasks, achieving an average F1 score of 0.84.
- Sensitivity Analysis:
- Evaluated model robustness under varying conditions such as distance, phone tilt angle, noise interference, and handheld methods.
- While environmental motion and finger dynamics had limited impact on overall performance, significant motion or device occlusion did reduce detection accuracy.
- Experimental Results:
-
Limitations and Future Directions:
- Sound Applicability: The current 16-19kHz signals may be slightly audible in certain cases, requiring further optimization of signal range or improved device frequency support (e.g., extension to iOS devices).
- Dynamic Scenarios: Calibration and optimization are needed for scenarios involving significant user movement (e.g., walking) or frequent device motion.
- Misclassification and Generalization: Accuracy declines for new users; future work should focus on building larger datasets for training to improve generalizability.
- Expression Accuracy Impact: Classification rates for specific expressions (e.g., sadness, surprise) are relatively low; further sensitivity analysis of AU (Action Unit) distribution is needed.
Conclusion
ExpresSense demonstrates the potential of acoustic-based facial expression detection without cameras, suitable for a range of real-time user emotion detection tasks. It provides theoretical and practical insights for privacy-friendly and lightweight next-generation smartphone applications.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can non-visual user facial expression detection be achieved using smartphone built-in speakers and microphones?Category: Social Agent Emotional and Nonverbal ExpressionSimilar questionsarrow_forward
- How can acoustic signal-based expression detection systems balance privacy protection and energy efficiency?Category: Social Agent Emotional and Nonverbal ExpressionSimilar questionsarrow_forward
- How effective is acoustic expression detection in complex scenarios such as occlusion and low light?Category: Social Agent Emotional and Nonverbal ExpressionSimilar questionsarrow_forward
Practical Problems
1- Expression detection is difficult for users in low light or masked scenarios, and traditional methods have poor privacy.Category: Social Agent Emotional and Nonverbal ExpressionSimilar questionsarrow_forward
- 100%
Quantified Canine: Inferring Dog Personality From Wearables
CHI '23· Human Pose & Activity Recognition +1
- 100%
HotFoot: Foot-Based User Identification using Thermal Imaging
CHI '23· Human Pose & Activity Recognition +1
- 100%
Enhancing Inertial Hand based HAR through Joint Representation of Language, Pose and Synthetic IMUs
UbiComp '24· Human Pose & Activity Recognition +1
- 100%
MLP-HAR: Boosting Performance and Efficiency of HAR Models on Edge Devices with Purely Fully Connected Layers
UbiComp '24· Human Pose & Activity Recognition +1
- 100%
Using Smartwatch Inertial Sensors to Recognize and Distinguish Between Car Drivers and Passengers
AutoUI '18· Human Pose & Activity Recognition +1
- 67%
Affective State Prediction from Smartphone Touch and Sensor Data in the Wild
CHI '22· Human Pose & Activity Recognition +2
- 67%
Sitting Posture Recognition and Feedback: A Literature Review
CHI '24· Human Pose & Activity Recognition +2
- 67%
PiaMuscle: Improving Piano Skill Acquisition by Cost-effectively Estimating and Visualizing Activities of Miniature Hand Muscles
CHI '25· Human Pose & Activity Recognition +1
- 67%
VAX: Using Existing Video and Audio-based Activity Recognition Models to Bootstrap Privacy-Sensitive Sensors
UbiComp '23· Human Pose & Activity Recognition +2
- 67%
PyroSense: 3D Posture Reconstruction Using Pyroelectric Infrared Sensing
UbiComp '24· Human Pose & Activity Recognition +2
Based on Jaccard similarity of research subtopics & professions (≥60%)