TouchScribe: Augmenting Non-Visual Hand-Object Interactions with Automated Live Visual Descriptions

Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)Vibrotactile Feedback & Skin StimulationBehavior Change & Reflection TechnologyAssistive Technology SpecialistsPhysicians, Nurses & Clinicians

Paper Title

TouchScribe: Augmenting Non-Visual Hand-Object Interactions with Automated Live Visual Descriptions

Publication Info

  • Topic area: Assistive technologies for blind and visually impaired individuals, focusing on real-time visual descriptions.
  • Keywords: Blind, visually impaired, hand-object interaction, visual descriptions, assistive technology, gesture recognition, real-time feedback, vision-language models, accessibility, object understanding.

Background and Problem

  • Problem / challenge: Blind and visually impaired (BLV) individuals face challenges accessing rich visual features of objects, such as color, text, and detailed patterns, which are inaccessible through touch alone. Current assistive technologies often rely on photo capturing and AI dialogue systems, which can be slow, inaccurate, and cumbersome.
  • Significance: Enabling BLV individuals to access detailed visual information in real-time enhances independence, object understanding, and everyday functionality, such as grocery shopping or categorizing items.
  • Motivation and related work: Previous systems like SeeingAI and VizLens provide limited visual feedback (e.g., text or color) and require explicit user input, such as photo capturing. Gesture-based systems have been explored but lack integration of rich, hierarchical descriptions. This paper builds on these gaps by leveraging natural hand-object interactions to provide live, adaptive visual descriptions.

Solution

  • Proposed approach: TouchScribe, a system that uses egocentric hand gestures to generate live, hierarchical visual descriptions of objects, including hand states, object labels, detailed descriptions, text, color, and comparisons.
  • Novelty:
    1. Integration of diverse hand gestures (e.g., hold, touch, point, swipe) to access hierarchical and adaptive visual descriptions.
    2. Use of a fine-tuned gesture recognition model combined with vision-language models (VLMs) for real-time object understanding.
    3. Hierarchical feedback design, enabling users to access brief, detailed, and comparative descriptions based on interaction context.
    4. Technical evaluation of gesture recognition accuracy, latency, and description accuracy in live settings.
  • Procedure and key techniques:
    • Gesture recognition using Google MediaPipe and a fine-tuned classification model for hand gestures and finger motions.
    • Keyframe extraction and object cropping using Hands23 to identify hand-object contact.
    • Description generation via Moondream and GPT-4o models for brief, detailed, and comparative feedback.
    • Real-time pipeline running on a neck-mounted smartphone setup with wide-angle camera coverage.

Results

  • Concrete findings:
    • Gesture recognition achieved an F1-score of 0.77, with highest accuracy for "hold" gestures (F1=0.84).
    • Description accuracy: 91.59% for brief object labels, 93.27% for detailed descriptions, 91.43% for comparative descriptions, and 67.83% for object texts.
    • Latency: Hand-state feedback averaged 0.56 seconds; brief descriptions took 5.36 seconds; detailed descriptions took 10.3 seconds; comparative descriptions took 14.0 seconds.
  • Advantage over baselines: TouchScribe provides richer and more integrated descriptions compared to prior systems, which typically focus on single information types (e.g., text or color). It also reduces reliance on photo capturing and turn-taking dialogues.
  • Experiments / evaluation:
    • User study with 8 BLV participants (ages 18–72) completing 4 object understanding tasks.
    • Tasks included understanding single objects, distinguishing between similar objects, sorting multiple objects, and selecting items based on specific criteria.
    • Participants rated TouchScribe as intuitive (M=5.63/7), accurate (M=5.5/7), and comprehensive (M=6.5/7), though noted moderate cognitive effort and a learning curve.
  • Limitations and future work:
    • Camera limitations: Wide-angle lens distortion affected gesture recognition accuracy.
    • Gesture recognition errors: False positives from unintentional hand movements.
    • Learning curve: Users required time to adapt to gesture mappings and hand positioning.
    • Future directions: Customizable gestures, improved gesture recognition using multimodal sensing, alternative camera configurations (e.g., smart glasses), and support for interactions beyond physical reach.

Summary

TouchScribe introduces a novel system for BLV individuals to access rich visual descriptions of objects through natural hand-object interactions. By leveraging egocentric gestures and vision-language models, it provides hierarchical feedback, including brief, detailed, and comparative descriptions. A user study demonstrated its effectiveness, with participants perceiving it as intuitive and accurate, though challenges like gesture recognition errors and camera limitations were noted. Future work aims to enhance gesture customization, multimodal sensing, and broader applicability in real-world contexts.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222163/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791308
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
7 authors
sell
Subtopics
Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille), Vibrotactile Feedback & Skin Stimulation, Behavior Change & Reflection Technology
work
Professions
Assistive Technology Specialists, Physicians, Nurses & Clinicians
article
Content Status
Full text indexed
hub
Related Papers
2 related papers