TouchScribe: Augmenting Non-Visual Hand-Object Interactions with Automated Live Visual Descriptions
Authors
Paper Title
TouchScribe: Augmenting Non-Visual Hand-Object Interactions with Automated Live Visual Descriptions
Publication Info
- Topic area: Assistive technologies for blind and visually impaired individuals, focusing on real-time visual descriptions.
- Keywords: Blind, visually impaired, hand-object interaction, visual descriptions, assistive technology, gesture recognition, real-time feedback, vision-language models, accessibility, object understanding.
Background and Problem
- Problem / challenge: Blind and visually impaired (BLV) individuals face challenges accessing rich visual features of objects, such as color, text, and detailed patterns, which are inaccessible through touch alone. Current assistive technologies often rely on photo capturing and AI dialogue systems, which can be slow, inaccurate, and cumbersome.
- Significance: Enabling BLV individuals to access detailed visual information in real-time enhances independence, object understanding, and everyday functionality, such as grocery shopping or categorizing items.
- Motivation and related work: Previous systems like SeeingAI and VizLens provide limited visual feedback (e.g., text or color) and require explicit user input, such as photo capturing. Gesture-based systems have been explored but lack integration of rich, hierarchical descriptions. This paper builds on these gaps by leveraging natural hand-object interactions to provide live, adaptive visual descriptions.
Solution
- Proposed approach: TouchScribe, a system that uses egocentric hand gestures to generate live, hierarchical visual descriptions of objects, including hand states, object labels, detailed descriptions, text, color, and comparisons.
- Novelty:
- Integration of diverse hand gestures (e.g., hold, touch, point, swipe) to access hierarchical and adaptive visual descriptions.
- Use of a fine-tuned gesture recognition model combined with vision-language models (VLMs) for real-time object understanding.
- Hierarchical feedback design, enabling users to access brief, detailed, and comparative descriptions based on interaction context.
- Technical evaluation of gesture recognition accuracy, latency, and description accuracy in live settings.
- Procedure and key techniques:
- Gesture recognition using Google MediaPipe and a fine-tuned classification model for hand gestures and finger motions.
- Keyframe extraction and object cropping using Hands23 to identify hand-object contact.
- Description generation via Moondream and GPT-4o models for brief, detailed, and comparative feedback.
- Real-time pipeline running on a neck-mounted smartphone setup with wide-angle camera coverage.
Results
- Concrete findings:
- Gesture recognition achieved an F1-score of 0.77, with highest accuracy for "hold" gestures (F1=0.84).
- Description accuracy: 91.59% for brief object labels, 93.27% for detailed descriptions, 91.43% for comparative descriptions, and 67.83% for object texts.
- Latency: Hand-state feedback averaged 0.56 seconds; brief descriptions took 5.36 seconds; detailed descriptions took 10.3 seconds; comparative descriptions took 14.0 seconds.
- Advantage over baselines: TouchScribe provides richer and more integrated descriptions compared to prior systems, which typically focus on single information types (e.g., text or color). It also reduces reliance on photo capturing and turn-taking dialogues.
- Experiments / evaluation:
- User study with 8 BLV participants (ages 18–72) completing 4 object understanding tasks.
- Tasks included understanding single objects, distinguishing between similar objects, sorting multiple objects, and selecting items based on specific criteria.
- Participants rated TouchScribe as intuitive (M=5.63/7), accurate (M=5.5/7), and comprehensive (M=6.5/7), though noted moderate cognitive effort and a learning curve.
- Limitations and future work:
- Camera limitations: Wide-angle lens distortion affected gesture recognition accuracy.
- Gesture recognition errors: False positives from unintentional hand movements.
- Learning curve: Users required time to adapt to gesture mappings and hand positioning.
- Future directions: Customizable gestures, improved gesture recognition using multimodal sensing, alternative camera configurations (e.g., smart glasses), and support for interactions beyond physical reach.
Summary
TouchScribe introduces a novel system for BLV individuals to access rich visual descriptions of objects through natural hand-object interactions. By leveraging egocentric gestures and vision-language models, it provides hierarchical feedback, including brief, detailed, and comparative descriptions. A user study demonstrated its effectiveness, with participants perceiving it as intuitive and accurate, though challenges like gesture recognition errors and camera limitations were noted. Future work aims to enhance gesture customization, multimodal sensing, and broader applicability in real-world contexts.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 67%
ElectroGrasp: Electrotactile Aids for Visually Impaired Individuals in Anticipatory Planning and Control of Grasp
CHI '26· Vibrotactile Feedback & Skin Stimulation +2
- 60%
Towards More Accessible Scientific PDFs for People with Visual Impairments: Step-by-Step PDF Remediation to Improve Tag Accuracy
CHI '25· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
Based on Jaccard similarity of research subtopics & professions (≥60%)