ImageExplorer: Multi-Layered Touch Exploration to Encourage Skepticism Towards Imperfect AI-Generated Image Captions
Authors
Title of the Paper
ImageExplorer: Multi-Layered Touch Exploration to Encourage Skepticism Towards Imperfect AI-Generated Image Captions
Paper Information
- Domain: Human-Computer Interaction, Artificial Intelligence, Accessibility Technology
- Keywords: Automatic Image Description, Alt Text, AI Error Skepticism, Touch Exploration, Screen Reader, Accessible Design, Visual Impairment, Image Understanding, Hierarchical Information Presentation, Human-Computer Interaction
Research Background and Issues
-
What problems or challenges did the authors identify?
- Current AI-generated image descriptions (e.g., alt-text) often lack accuracy, with frequent omissions or errors.
- Visually impaired users tend to overly trust these AI-generated descriptions, especially when other supporting information is absent.
- Existing image exploration tools fail to support users in verifying AI-generated descriptions.
-
Why is this issue important?
- Visually impaired individuals heavily rely on alt-text or descriptions to understand online image content. Erroneous descriptions can lead to misinterpretation of information.
- Providing rich and accurate methods for exploring information can enhance visually impaired users' understanding of images and reduce blind trust in flawed AI-generated descriptions.
-
Research Motivation and Related Work
- The authors observed that tactile exploration and hierarchical content presentation help visually impaired users better understand images, but these methods have not been integrated into a single system.
- Additionally, the authors identified limitations in existing systems like Facebook and Seeing AI, prompting the development of a new approach that combines tactile exploration with hierarchical presentation.
Solution
-
What methods or solutions did the authors propose?
- Developed a new tactile exploration system, "ImageExplorer," which integrates hierarchical information presentation with tactile interaction to provide visually impaired users with a detailed and structured image exploration experience.
- Designed the system with a two-layer structure:
- Layer 1: Presents the main objects in the image and their boundaries.
- Layer 2: Provides finer-grained information about sub-objects, such as specific attributes and relationships.
-
What are the innovative aspects of this solution?
- Integrated multiple deep learning technologies (e.g., Mask R-CNN, Google Cloud Vision, DenseCap) to ensure comprehensive information coverage.
- Combined touch interaction with hierarchical information, supporting spatial relationship exploration while allowing users to control detailed information retrieval.
- Enhanced user experience by enabling layer switching via double-tap, improving user autonomy and interaction flexibility.
-
What are the implementation steps and key technologies used?
- Used deep learning models to extract image content and scene hierarchy.
- Built an iOS application supporting tactile interaction with audio feedback.
- Defined hierarchical display rules to ensure orderly and accessible information organization.
- Provided features like object boundary rendering, audio prompts, double-tap for detailed exploration, and progress feedback during element exploration.
Research Outcomes
-
What specific outcomes were achieved?
- Users became more skeptical of AI-generated descriptions after using ImageExplorer, especially when descriptions were partially or entirely incorrect.
- Compared to Seeing AI and Facebook, ImageExplorer helped users interpret image content more accurately and identify errors in descriptions.
-
How does it compare to existing solutions?
- Offers more detailed and structured information than Facebook and Seeing AI.
- Multi-layer information presentation allows users to access more detailed content on demand, whereas single-layer systems (e.g., Seeing AI) provide limited information.
- Beyond detail, ImageExplorer retains spatial relationship information, aiding users in constructing mental images.
-
What were the experimental or evaluation results?
- In an experiment with 12 visually impaired participants:
- Multi-layer exploration (ImageExplorer) led participants to produce more accurate interpretations of erroneous descriptions.
- In scoring, tactile systems (e.g., ImageExplorer) significantly reduced user ratings of flawed descriptions compared to text-based systems (e.g., Facebook).
- Users generally considered ImageExplorer to provide the most comprehensive information, though text-based systems (e.g., Facebook) were still favored for ease of use.
- In an experiment with 12 visually impaired participants:
-
Limitations and Future Directions
- Limitations:
- Tactile exploration requires significantly more time than text-based exploration, potentially causing fatigue.
- Despite using a multi-model approach, ImageExplorer occasionally provides inaccurate or incomplete information.
- Future Directions:
- Develop smarter structured description generation models to further improve the accuracy and comprehensiveness of information.
- Optimize interaction feedback mechanisms (e.g., touch-and-hold as an alternative to double-tap) to reduce cognitive load for users.
- Explore integration of text and touch features to create more flexible system designs that cater to diverse user needs.
- Limitations:
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can blind and low-vision users' blind trust in flawed AI-generated image descriptions be reduced?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- Can haptic exploration and layered information presentation help blind users more accurately understand image content and identify description errors?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- How does multi-level haptic interaction affect users' skepticism toward AI errors and usage experience?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
Practical Problems
1- Blind users easily blindly trust flawed AI-generated image descriptions and struggle to accurately understand content.Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- 67%
Towards Enabling Blind People to Independently Write on Printed Forms
CHI '19· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 67%
Examining Visual Semantic Understanding in Blind and Low-Vision Technology Users
CHI '21· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 67%
TableNarrator: Making Image Tables Accessible to Blind and Low Vision People
CHI '25· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 67%
Scene Text Access: A Comparison of Mobile OCR Modalities for Blind Users
IUI '19· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 60%
Designing Accessible Obfuscation Support for Blind Individuals’ Visual Privacy Management
CHI '24· Explainable AI (XAI) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)