Leveraging Biometric-Rich Hand Gestures for Head-Mounted Display Authentication
Honorable MentionAuthors
Paper Title
Leveraging Biometric-Rich Hand Gestures for Head-Mounted Display Authentication
Publication Info
- Topic area: Behavioral biometric authentication for head-mounted displays (HMDs).
- Keywords: HMD authentication, behavioral biometrics, free-form gestures, hand joint motion, anomaly detection, 3D gestures, observation attacks, usability, longitudinal study, CNN-LSTM autoencoder.
Background and Problem
- Problem / challenge: Traditional HMD authentication methods (e.g., passwords) are cumbersome, prone to shoulder-surfing attacks, and lack usability. Existing behavioral biometric approaches rely on limited motion data (e.g., controllers) and predefined gestures, which restrict expressiveness and security evaluations.
- Significance: Securing HMDs is critical as they store sensitive data and are increasingly used in diverse applications like gaming, education, and healthcare. A robust, usable, and secure authentication method is needed to address real-world attack scenarios.
- Motivation and related work: Prior work demonstrated the feasibility of motion-based biometrics but faced limitations in gesture diversity, evaluation against skilled attackers, and longitudinal performance. This paper builds on these findings to address these gaps by integrating free-form gestures with hand joint biometrics.
Solution
- Proposed approach: A knowledge-driven behavioral biometric authentication system combining user-defined 3D gestures with hand joint motion signals, processed using a CNN-LSTM autoencoder for anomaly detection.
- Novelty:
- Integration of free-form 3D gestures as knowledge-based credentials with hand joint biometrics for enhanced security.
- Comprehensive evaluation against worst-case observation attacks using skilled attackers and longitudinal usability over one week.
- Development of a publicly available gesture-entry system and dataset for reproducibility.
- Insights into the impact of gesture complexity and joint signal contributions on authentication performance.
- Procedure and key techniques:
- Users create gestures by pinching their index finger and thumb and drawing in 3D space.
- Hand joint motion signals (position and rotation) are captured and preprocessed (e.g., smoothing, normalization).
- A hybrid CNN-LSTM autoencoder is trained to detect anomalies in gesture patterns.
- The system is evaluated through user studies, including observation attack simulations and multi-session recall tests.
Results
- Concrete findings:
- Dictionary gestures achieved low error rates (EER = 5.08% for 3D, 5.68% for 2D).
- User-generated gestures resisted observation attacks with EERs of 3.58% (baseline) and 4.55% (3D policy).
- Longitudinal recall performance improved with retraining, reducing EERs to 9.82% (baseline) and 11.32% (3D policy) after one week.
- Advantage over baselines:
- Improved security compared to prior motion-based systems (e.g., EER = 12.4% in earlier work).
- Stronger defense against skilled observation attacks due to richer biometric signals (237 features).
- Comparable usability and recall rates for both baseline and 3D-enforced systems.
- Experiments / evaluation:
- User study with 20 participants, capturing 21,600 gestures across two conditions (baseline and 3D policy).
- Observation attack study with 10 attackers simulating worst-case scenarios.
- Longitudinal recall tests over one week with retraining to address behavioral variability.
- Limitations and future work:
- Small and homogeneous participant sample; future studies should include diverse demographics.
- Single-device platform (Meta Quest 3) and fixed sampling rate (60 Hz); additional devices and higher rates may improve results.
- Limited exploration of multimodal data (e.g., gaze, head motion) and fine-tuning of classifiers.
- Observation attack study used a single-view camera; multi-perspective setups could enhance evaluations.
Summary
This paper presents a behavioral biometric authentication system for HMDs that combines user-defined 3D gestures with hand joint motion signals. The system demonstrated strong security against skilled observation attacks (EER = 3.58%–4.55%) and maintained usability and memorability over one week, with retraining reducing EERs to 9.82%. By leveraging a hybrid CNN-LSTM autoencoder and a rich set of biometric features, the approach outperformed prior motion-based systems. Future work should expand participant diversity, explore multimodal data, and refine classifiers to further enhance real-world applicability.
Research Questions / Practical Problems
Question signals indexed for this paper.
Based on Jaccard similarity of research subtopics & professions (≥60%)