Interface Design for Crowdsourcing Hierarchical Multi-Label Text Annotations

Crowdsourcing Task Design & Quality ControlField StudiesHCI ResearchersAmazon Mechanical Turk Workers

Title of the Paper

Interface Design for Crowdsourcing Hierarchical Multi-Label Text Annotations

Paper Information

  • Domain: Human-Computer Interaction (HCI), Crowdsourced Data Annotation, Machine Learning
  • Keywords: Crowdsourcing, Text Annotation, User Experience Design, Hierarchy, Multi-Label Classification, Classification Interface, User Interface Design, Annotation Efficiency, Text Processing, AI Annotation
  • Conference: CHI 2023 (ACM Conference on Human Factors in Computing Systems)

Research Background and Problem

  1. Identified Challenges:

    • Generating supervised learning datasets through manual annotation is not only expensive but also faces issues of subjectivity and task complexity. For multi-label tasks, especially hierarchical label classification tasks, how to improve annotation quality and efficiency remains unclear.
    • Highly subjective annotation tasks (e.g., annotating vaccine-related misinformation) further highlight challenges in accuracy and consistency.
    • There is no clear answer on how to design annotation task interfaces based on conceptual hierarchies to enhance performance (e.g., grouping similar concepts, reducing cognitive load).
  2. Significance:

    • The high cost and demand for high-quality data annotation are critical for the performance of machine learning models.
    • Annotation challenges in multi-label tasks have broad practical applications (e.g., misinformation research, social media analysis).
  3. Research Motivation:

    • To explore how integrating hierarchical information into user interface design can improve the quality and efficiency of annotation tasks.
    • To propose methods tailored to different task scenarios and provide quantitative evaluations on how to best utilize limited budgets to collect high-quality data.

Solution

Methods and Innovations

  1. Introduction of Multiple Annotation Interface Designs:

    • Comparison of three interface layouts:
      • Binary-Label: Annotating one "yes/no" question at a time.
      • Flat Multi-Label: Displaying all labels in a flat layout, allowing multiple selections.
      • Hierarchical Multi-Label: Labels are organized hierarchically, allowing step-by-step selection.
    • Integration of three annotation logics:
      • Single-Pass, Multi-Pass, Hierarchical Multi-Pass.
  2. Implementation of Key Experimental Settings:

    • Fixed budget, balancing the cognitive load of workers and the trade-offs of multi-worker collaboration.
    • Comparison of the effects of randomly grouped labels versus hierarchy-guided grouped labels.
  3. Core Techniques and Processes:

    • Hierarchical Annotation Design:
      • Development of a hierarchical labeling system for vaccine misinformation (e.g., top-level label "Health Risks" with sub-labels like "Specific Side Effects").
    • Annotation Workflow:
      • Workers undergo customized training and assessment (e.g., tutorials and tests) before starting tasks.
      • A custom crowdsourcing platform is used to control task display logic and collect multi-round annotation data.

Research Findings

  1. Key Conclusions:

    • Annotation schemes incorporating hierarchical information improved workers' F1 scores (+0.16 points), particularly in:
      • Grouping Similar Concepts significantly enhanced annotation quality (F1 = 0.50 for hierarchy-based grouping versus 0.34 for random grouping).
      • Significant performance improvement for difficult annotation examples (+0.40 F1 score).
      • Filtering irrelevant negative examples (pre-detection) improved annotation precision, increasing F1 from 0.50 to 0.57.
  2. Comparison with Existing Methods:

    • The single-pass hierarchical multi-label scheme combined with a majority vote mechanism achieved the best performance within a fixed budget (F1 = 0.70).
    • Multi-pass schemes (especially hierarchical progressive schemes) showed advantages in recall but were slightly inferior in precision.
  3. Experimental and Evaluation Results:

    • Detailed comparisons of different interface designs and annotation logics:
      • Hierarchical grouping outperformed random grouping.
      • Hierarchical multi-label schemes were beneficial for complex and highly subjective tasks.
      • Overall analysis indicated that annotator performance was significantly influenced by annotation frequency, true positive frequency, and task grouping.
  4. Limitations and Future Directions:

    • Limitations:
      • Experimental results did not include "Binary-Label" schemes.
      • The generalizability of hierarchical structures remains uncertain (requires validation on more types of hierarchical tasks).
      • Testing was limited to the Amazon Mechanical Turk platform, and results may not generalize to other crowdsourcing platforms.
    • Future Directions:
      • Further optimization of hierarchical label design.
      • Exploration of applicability to larger budgets, diverse tasks, and more complex hierarchical labels.
      • Enhanced exploration of sub-task decomposition and topic pre-filtering mechanisms.

Conclusion

  • This study systematically analyzed the impact of integrating hierarchical information into interface design on annotation quality and efficiency in highly subjective text annotation tasks.
  • The findings provide valuable guidance for building complex data annotation tools on future crowdsourcing platforms.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/95763/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581431
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Crowdsourcing Task Design & Quality Control, Field Studies
work
Professions
HCI Researchers, Amazon Mechanical Turk Workers
article
Content Status
Full text indexed
hub
Related Papers
8 related papers