Crowdsourcing Thumbnail Captions Using Time-Constrained Methods

Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille)Crowdsourcing Task Design & Quality ControlAssistive Technology SpecialistsAmazon Mechanical Turk Workers

Document Title

Crowdsourcing Thumbnail Captions via Time-Constrained Methods

Document Information

  • Subject Area: User interface and human-computer interaction, specifically designing accessible image descriptions for visually impaired users
  • Keywords: Image captions, crowdsourcing, accessibility, annotation interface, time-constrained methods, visually impaired users, artificial intelligence, language processing

Research Background and Problem

  • Problems and Challenges:
    • Existing image captioning methods typically provide single-level descriptions, which fail to meet the needs of different usage scenarios.
    • For visually impaired users, short captions (i.e., thumbnail captions) facilitate quick content browsing, but current "alternative text" is often overly verbose or insufficiently detailed.
    • The quality of automated image captioning tools (AAT) remains inadequate, while human-generated captions are costly and slow.
  • Significance:
    • With the rapid growth of image data on the internet, providing robust image captioning mechanisms is critical for accessible information dissemination.
    • In accessibility contexts, thumbnail captions can significantly enhance browsing efficiency for visually impaired users.
  • Research Motivation and Related Work:
    • Inspired by visually impaired users' need for progressively detailed interactive captions, the study explores systematic collection of image captions with varying levels of detail.
    • Current methods for evaluating image descriptions or captions (e.g., BLEU, ROUGE) have limitations, necessitating the exploration of more effective model metrics.

Solution

Method and Framework

  • The authors propose a time-constrained crowdsourcing method combined with other text description approaches:
    1. Control: Standard descriptions based on the MS COCO image caption dataset.
    2. Comprehensive: Instructing online workers to generate more detailed captions.
    3. Essential: Instructing workers to describe only the most important information in the image.
    4. Timed (Time-Constrained Method): Limiting image viewing time to 500 milliseconds, assuming this method encourages recall of the most critical image details.

Implementation Steps and Key Techniques

  1. Crowdsourcing Collection Experiment:
    • Online workers were recruited from Amazon Mechanical Turk to generate captions using the four methods mentioned above.
    • Image data was sourced from the MS COCO dataset, covering themes such as events/scenes, people, and objects.
  2. Evaluation Dimensions:
    • Captions were evaluated for correctness, fluency, and level of detail using both human and automated metrics.
    • Human evaluation followed existing frameworks, quantifying correctness, fluency, and detail on a 0-100 scale.
  3. Model-Based Evaluation:
    • Correctness Metrics: SPICE_f and ViLBERTScore_f were used to assess caption correctness.
    • Detail Metrics: Caption detail was quantified using noun phrase counts (NPs) and cross-entropy.

Innovations

  • Proposed a novel time-constrained method for generating thumbnail captions, leveraging human rapid visual processing capabilities.
  • Systematically explored how traditional textual prompts or time constraints can differentiate levels of image description.
  • Validated the effectiveness of existing model metrics for evaluating multi-level content, providing a foundation for future automated evaluation systems.

Research Findings

Specific Results

  1. Effectiveness Validation:
    • The time-constrained (Timed) method significantly generated more concise captions while maintaining comparable correctness and fluency to other methods.
    • Human evaluation showed that captions generated by the Timed method were significantly less detailed than the control group, while textual prompt methods showed no significant differences.
  2. Consistency and Efficiency:
    • Captions generated using the Timed method demonstrated high consistency across multiple experiments.
    • The Timed method was more efficient, with workers spending significantly less time generating captions compared to other methods.
  3. Automated Metric Evaluation:
    • SPICE_f and ViLBERTScore_f showed moderate positive correlations with human correctness scores (correlation coefficients of 0.29 and 0.18, respectively).
    • Cross-entropy and noun phrase counts effectively reflected caption detail levels, showing significant correlations with human evaluations.

Advantages Over Existing Solutions

  • Compared to traditional textual prompt mechanisms, the time-constrained method automatically guides workers to generate thumbnail captions without explicit instructions.
  • The method is more efficient, offering better cost and time control for online crowdsourcing.
  • Provides a systematic evaluation mechanism, laying the groundwork for future user experience optimization.

Limitations and Future Directions

  • Contextual Applicability of Captions:
    • The study did not evaluate the effectiveness of generated captions in real-world usage scenarios for visually impaired users.
    • Future research should test the practical usability and acceptability of these captions for blind users' browsing efficiency.
  • Diversity in Levels of Detail:
    • The study primarily succeeded in collecting thumbnail captions but did not significantly expand levels of detailed descriptions.
    • Future work should explore more refined task instructions or dynamic feedback mechanisms to collect multi-level descriptions.
  • Improvements in Model-Based Evaluation:
    • Current metrics (especially ViLBERTScore_f) exhibit biases toward caption length, necessitating the development of more universal evaluation models.

Conclusion

This study validates the effectiveness of the time-constrained method for generating accessible thumbnail captions. It highlights the limitations of existing textual prompt methods in producing content with varying levels of detail and offers novel insights and directions for large-scale automated dataset construction in the future.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/iui/79942/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3490099.3511136
At a Glance

Paper Snapshot

fact_check
dataset
Source
IUI
calendar_month
Year
2022
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Visual Impairment Technologies (Screen Readers, Tactile Graphics, Braille), Crowdsourcing Task Design & Quality Control
work
Professions
Assistive Technology Specialists, Amazon Mechanical Turk Workers
article
Content Status
Full text indexed
hub
Related Papers
2 related papers