HandyTrak: Recognizing the Holding Hand on a Commodity Smartphone from Body Silhouette Images
Authors
Document Title
HandyTrak: Recognizing the Holding Hand on a Commodity Smartphone from Body Silhouette Images
Document Information
- Subject Area: Human-Computer Interaction, Mobile Computing, Computer Vision
- Keywords: Mobile Computing, Computer Vision, Commodity Smartphone, Deep Learning, Hand Pattern Recognition, User Interface, Real-Time Tracking, Body Silhouette, AI Applications
Research Background and Problem
-
Identified Problems or Challenges:
- The way users hold their smartphones directly impacts their user experience, especially when one-handed operation is difficult. As smartphone screen sizes increase, it becomes harder for users to reach certain areas of the screen with their thumbs.
- Existing methods (e.g., Apple's Reachability feature) require manual activation by the user or can only recognize hand patterns in limited scenarios (e.g., during unlocking), failing to achieve continuous tracking.
- Some methods also require additional hardware, making them unsuitable for commodity smartphones.
-
Significance:
- Improving user experience: Automatically adapting UI layouts to reduce operational difficulty.
- Enhancing the efficiency of intelligent interaction, reducing the need for users to manually adjust the UI.
-
Research Motivation and Related Work:
- Existing studies infer hand patterns using input data such as touch events or inertial measurement unit (IMU) sensors, but these require explicit user interaction or rely on additional hardware.
- The goal of this study is to develop a novel method that leverages the front-facing camera of smartphones to capture users' body silhouettes, enabling continuous tracking without user input.
Solution
-
Proposed Solution: HandyTrak is an AI-based software system that uses the built-in front-facing camera of commodity smartphones to capture body silhouette images. Through deep learning algorithms, it continuously identifies whether the user is holding the phone with their left hand, right hand, or both hands.
-
Innovations:
- No additional hardware is required, relying solely on the smartphone's built-in front-facing camera.
- Achieves continuous hand pattern tracking without explicit user interaction.
- Proposes a deep learning method for classification using body silhouette images.
-
Implementation Steps and Techniques:
- Image Preprocessing:
- Normalize the size of images captured by the front-facing camera (224×224) and perform human segmentation using the FCN-ResNet101 model.
- Deep Learning:
- Design a custom classification network (HandyNet) based on the pre-trained VGG16 model.
- Classify the segmented silhouette images into left-hand, right-hand, or both-hand categories, with results output through a softmax layer.
- System Deployment:
- HandyTrak is trained and tested in various scenarios, with model parameters fine-tuned through experimentation.
- Image Preprocessing:
Research Outcomes
-
Specific Results:
- HandyTrak achieved an average classification accuracy of 89.03% when users held the phone (sampling rate: 2Hz).
- For tilted phone-holding scenarios, classification accuracy was higher (95.84%, standard deviation: 5.05%).
- HandyTrak achieved 92.3% classification accuracy within a one-second sliding window and performed consistently across different front-facing camera positions.
-
Advantages:
- Compared to existing methods requiring explicit input or additional hardware, HandyTrak is more convenient and applicable.
- It can dynamically adjust UI layouts in smartphone applications, such as adapting virtual button positions during selfies.
-
Experimental or Evaluation Results:
- The system performed consistently across three common scenarios (standing, sitting, and tabletop support).
- Classification accuracy significantly improved in tilted phone-holding scenarios.
- Task-specific evaluations showed the following accuracies: 87.56% during unlocking, 88.37% during in-app browsing, and 92.68% during selfies.
- Further studies indicated that the system maintained an accuracy of 88.36% even when users were walking.
-
Limitations and Future Directions:
- Coverage: Currently, the system is trained only for three usage postures (standing, sitting, and tabletop support). Future work should expand to more complex scenarios (e.g., lying down, walking).
- User Independence: This study primarily uses user-dependent training models. User-independent models showed weaker performance and require more data for large-scale training.
- Energy Consumption and Privacy: Continuous camera usage increases power consumption. Future work could optimize system activation by integrating IMU sensors.
- Low-Light Adaptability: The system performs poorly in low-light conditions. Future work could incorporate depth cameras to address nighttime usage.
- Further Improvements: Enhanced image segmentation techniques and multimodal data (e.g., depth images) could improve model performance.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- Can a standard smartphone front camera detect users' phone-holding mode in real time using only body silhouette images?Category: Input Recognition, Touch Modeling, and Interaction Decoding MethodsSimilar questionsarrow_forward
- How can deep learning identify whether users hold phones with left hand, right hand, or both hands?Category: Input Recognition, Touch Modeling, and Interaction Decoding MethodsSimilar questionsarrow_forward
- How can hand-mode recognition accuracy and applicability be improved without additional hardware?Category: Input Recognition, Touch Modeling, and Interaction Decoding MethodsSimilar questionsarrow_forward
Practical Problems
1- Large-screen phones are difficult to operate, especially for one-handed touch.Category: Input Recognition, Touch Modeling, and Interaction Decoding MethodsSimilar questionsarrow_forward
- 100%
HandSee: Enabling Full Hand Interaction on Smartphone with Front Camera-based Stereo Vision
CHI '19· Hand Gesture Recognition +1
- 75%
M3 Gesture Menu: Design and Experimental Analyses of Marking Menus for Touchscreen Mobile Interaction
CHI '18· Hand Gesture Recognition
- 75%
PinchList: Leveraging Pinch Gestures for Hierarchical List Navigation on Smartphones
CHI '19· Hand Gesture Recognition
- 67%
EyeEcho: Continuous and Low-power Facial Expression Tracking on Glasses
CHI '24· Hand Gesture Recognition +3
- 60%
GestAKey: Touch Interaction on Individual Keycaps
CHI '18· Hand Gesture Recognition +1
- 60%
Characterizing Finger Pitch and Roll Orientation During Atomic Touch Actions
CHI '18· Hand Gesture Recognition +1
- 60%
VirtualGrasp: Leveraging Experience of Interacting with Physical Objects to Facilitate Digital Object Retrieval
CHI '18· Hand Gesture Recognition +1
- 60%
Projective Windows: Bringing Windows in Space to the Fingertip
CHI '18· Hand Gesture Recognition +1
- 60%
Pinpointing: Precise Head- and Eye-Based Target Selection for Augmented Reality
CHI '18· Eye Tracking & Gaze Interaction +1
- 60%
Training Person-Specific Gaze Estimators from User Interactions with Multiple Devices
CHI '18· Eye Tracking & Gaze Interaction +1
Based on Jaccard similarity of research subtopics & professions (≥60%)