Ga11y: an Automated GIF Annotation System for Visually Impaired Users
Authors
Document Title
Ga11y: An Automated GIF Annotation System for Visually Impaired Users
Document Information
- Subject Area: Human-Computer Interaction, Assistive Technology, Accessibility Design
- Keywords: GIF, Image Description, Blind, Low Vision, Text Description, Human Annotation, Crowdsourcing, Accessibility
Research Background and Problem
- Problem and Challenges: Dynamic GIF images are widely used on the internet and social media but rarely include textual descriptions, making it difficult for blind or low vision (BLV) users to understand their content. Existing computer vision technologies face limitations in describing the dynamic and ambiguous nature of GIF content.
- Importance: GIFs are rich carriers of emotional expression and cultural context. For BLV users, the inaccessibility of this information hinders their ability to equally participate in online communication.
- Research Motivation and Related Work: Previous efforts have successfully addressed static image descriptions, such as using alt text to describe image content. However, annotating dynamic GIFs remains a highly challenging problem, and existing research and industrial practices have shown limited effectiveness in helping BLV users comprehend GIF content.
Solution
- Method or Solution: This study introduces the Ga11y system, an automated GIF annotation system that combines machine intelligence with crowdsourcing.
- Innovations:
- Integration of an Android client, backend server, and web-based human annotation interface.
- Adoption of a "semi-structured template" design for annotation tasks to significantly improve annotation quality.
- Implementation of user-interactive GIF reconstruction and similarity matching algorithms.
- Implementation Steps and Key Technologies:
- Develop an Android client to monitor GIF elements on the user's screen, provide an interactive interface for annotation requests, and record screen GIFs.
- Parse and store keyframes on the server while matching GIFs in the database.
- If no similar GIF exists in the database, use computer vision services (Google Vision, Microsoft Azure) to generate machine annotations.
- Provide a web platform for volunteers to write detailed human annotations based on GIF content.
- Use semi-structured task design combining free expression and guided prompts to enhance annotation quality.
- Synchronize descriptions to the server to support future rapid matching.
Research Outcomes
- Specific Results:
- The Ga11y system efficiently annotates dynamic GIFs, allowing BLV users to listen to textual descriptions via the Android client.
- The system achieved a high System Usability Scale (SUS) score of 89.1/100, with users rating it as "highly user-friendly."
- Machine-generated text provides basic, timely information for unannotated GIFs, while human annotations deliver richer and higher-quality descriptions.
- Advantages Over Existing Solutions:
- Integration of human and machine intelligence enables sustainable updates to the annotation database.
- Semi-structured annotation style is perceived by users as clearer and more informative than freeform or structured descriptions.
- Open-source design supports further development and expansion.
- Experiment and Evaluation Results:
- Through two-stage user studies and testing with over 12 participants, the system demonstrated excellent performance in improving BLV users' online communication quality.
- During a 3-day real-world usage period, the system processed 548 annotation requests, showcasing its potential for high-frequency use.
- Limitations and Future Directions:
- Overly detailed annotations need improvement, such as allowing users to adjust the level of detail in descriptions.
- Current data processing involves privacy concerns, which could be addressed by transitioning to localized storage and recognition.
- Users expressed a desire for support across more languages and platforms (e.g., iOS).
- Long-term data collection and analysis require further validation, including testing with more diverse and global user groups.
Conclusion
Ga11y combines machine learning and human collaboration to address the technical and interaction challenges of dynamic GIF annotation. This research supports BLV users in better understanding and participating in internet culture and provides critical technical references for future development of GIF platforms.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- Can GIF content be described at high quality through a combination of automation and human annotation?Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
- Can semi-structured templates improve description quality in GIF annotation tasks?Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
- How can GIF annotation systems improve visually impaired users' ability to understand animated content?Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
Practical Problems
1- Visually impaired users cannot understand animated GIFs, limiting their participation in online social and cultural exchange.Category: Chart, Image, and Visual Content AccessibilitySimilar questionsarrow_forward
- 100%
Understanding Blind Screen-Reader Users' Experiences of Digital Artboards
CHI '21· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 100%
What makes web data tables accessible? Insights and a tool for rendering accessible tables for people with visual impairments
CHI '22· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 100%
Modeling Touch-based Menu Selection Performance of Blind Users via Reinforcement Learning
CHI '23· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 100%
"It Brought Me Joy": Opportunities for Spatial Browsing in Desktop Screen Readers
CHI '25· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 100%
InSupport: Proxy Interface for Enabling Efficient Non-Visual Interaction with Web Data Records
IUI '22· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 100%
OmniScribe: Authoring Immersive Audio Descriptions for 360° Videos
UIST '22· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 100%
MagnePins: A Modular, Affordable, and DIY Refreshable Braille and Tactile Display
UIST '25· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 67%
A Face Recognition Application for People with Visual Impairments: Understanding Use Beyond the Lab
CHI '18· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 67%
Enabling People with Visual Impairments to Navigate Virtual Reality with a Haptic and Auditory Cane Simulation
CHI '18· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
- 67%
SteeringWheel: A Locality-Preserving Magnification Interface for Low Vision Web Browsing
CHI '18· Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)
Based on Jaccard similarity of research subtopics & professions (≥60%)