Interface Design for Crowdsourcing Hierarchical Multi-Label Text Annotations
Title of the Paper
Interface Design for Crowdsourcing Hierarchical Multi-Label Text Annotations
Paper Information
- Domain: Human-Computer Interaction (HCI), Crowdsourced Data Annotation, Machine Learning
- Keywords: Crowdsourcing, Text Annotation, User Experience Design, Hierarchy, Multi-Label Classification, Classification Interface, User Interface Design, Annotation Efficiency, Text Processing, AI Annotation
- Conference: CHI 2023 (ACM Conference on Human Factors in Computing Systems)
Research Background and Problem
-
Identified Challenges:
- Generating supervised learning datasets through manual annotation is not only expensive but also faces issues of subjectivity and task complexity. For multi-label tasks, especially hierarchical label classification tasks, how to improve annotation quality and efficiency remains unclear.
- Highly subjective annotation tasks (e.g., annotating vaccine-related misinformation) further highlight challenges in accuracy and consistency.
- There is no clear answer on how to design annotation task interfaces based on conceptual hierarchies to enhance performance (e.g., grouping similar concepts, reducing cognitive load).
-
Significance:
- The high cost and demand for high-quality data annotation are critical for the performance of machine learning models.
- Annotation challenges in multi-label tasks have broad practical applications (e.g., misinformation research, social media analysis).
-
Research Motivation:
- To explore how integrating hierarchical information into user interface design can improve the quality and efficiency of annotation tasks.
- To propose methods tailored to different task scenarios and provide quantitative evaluations on how to best utilize limited budgets to collect high-quality data.
Solution
Methods and Innovations
-
Introduction of Multiple Annotation Interface Designs:
- Comparison of three interface layouts:
- Binary-Label: Annotating one "yes/no" question at a time.
- Flat Multi-Label: Displaying all labels in a flat layout, allowing multiple selections.
- Hierarchical Multi-Label: Labels are organized hierarchically, allowing step-by-step selection.
- Integration of three annotation logics:
- Single-Pass, Multi-Pass, Hierarchical Multi-Pass.
- Comparison of three interface layouts:
-
Implementation of Key Experimental Settings:
- Fixed budget, balancing the cognitive load of workers and the trade-offs of multi-worker collaboration.
- Comparison of the effects of randomly grouped labels versus hierarchy-guided grouped labels.
-
Core Techniques and Processes:
- Hierarchical Annotation Design:
- Development of a hierarchical labeling system for vaccine misinformation (e.g., top-level label "Health Risks" with sub-labels like "Specific Side Effects").
- Annotation Workflow:
- Workers undergo customized training and assessment (e.g., tutorials and tests) before starting tasks.
- A custom crowdsourcing platform is used to control task display logic and collect multi-round annotation data.
- Hierarchical Annotation Design:
Research Findings
-
Key Conclusions:
- Annotation schemes incorporating hierarchical information improved workers' F1 scores (+0.16 points), particularly in:
- Grouping Similar Concepts significantly enhanced annotation quality (F1 = 0.50 for hierarchy-based grouping versus 0.34 for random grouping).
- Significant performance improvement for difficult annotation examples (+0.40 F1 score).
- Filtering irrelevant negative examples (pre-detection) improved annotation precision, increasing F1 from 0.50 to 0.57.
- Annotation schemes incorporating hierarchical information improved workers' F1 scores (+0.16 points), particularly in:
-
Comparison with Existing Methods:
- The single-pass hierarchical multi-label scheme combined with a majority vote mechanism achieved the best performance within a fixed budget (F1 = 0.70).
- Multi-pass schemes (especially hierarchical progressive schemes) showed advantages in recall but were slightly inferior in precision.
-
Experimental and Evaluation Results:
- Detailed comparisons of different interface designs and annotation logics:
- Hierarchical grouping outperformed random grouping.
- Hierarchical multi-label schemes were beneficial for complex and highly subjective tasks.
- Overall analysis indicated that annotator performance was significantly influenced by annotation frequency, true positive frequency, and task grouping.
- Detailed comparisons of different interface designs and annotation logics:
-
Limitations and Future Directions:
- Limitations:
- Experimental results did not include "Binary-Label" schemes.
- The generalizability of hierarchical structures remains uncertain (requires validation on more types of hierarchical tasks).
- Testing was limited to the Amazon Mechanical Turk platform, and results may not generalize to other crowdsourcing platforms.
- Future Directions:
- Further optimization of hierarchical label design.
- Exploration of applicability to larger budgets, diverse tasks, and more complex hierarchical labels.
- Enhanced exploration of sub-task decomposition and topic pre-filtering mechanisms.
- Limitations:
Conclusion
- This study systematically analyzed the impact of integrating hierarchical information into interface design on annotation quality and efficiency in highly subjective text annotation tasks.
- The findings provide valuable guidance for building complex data annotation tools on future crowdsourcing platforms.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can interface design integrate conceptual hierarchy information to improve quality and efficiency in multi-label text annotation tasks?Category: Human-in-the-Loop Labeling and Example SelectionSimilar questionsarrow_forward
- Does hierarchical multi-label annotation have significant advantages in complex and highly subjective tasks?Category: Human-in-the-Loop Labeling and Example SelectionSimilar questionsarrow_forward
- Compared with random grouping, can hierarchy-based grouping significantly improve annotation accuracy and consistency?Category: Human-in-the-Loop Labeling and Example SelectionSimilar questionsarrow_forward
Practical Problems
1- Text annotation tasks are complex, costly, and prone to subjective inconsistency.Category: Human-in-the-Loop Labeling and Example SelectionSimilar questionsarrow_forward
- 75%
Online Sequencing of Non-Decomposable Macrotasks in Expert Crowdsourcing
CHI '18· Crowdsourcing Task Design & Quality Control
- 60%
Crowdlicit: A System for Conducting Distributed End-User Elicitation and Identification Studies
CHI '19· Crowdsourcing Task Design & Quality Control +1
- 60%
Crowdsourced Detection of Emotionally Manipulative Language
CHI '20· AI Ethics, Fairness & Accountability +1
- 60%
Can we crowdsource Tacton similarity perception and metaphor ratings?
CHI '23· Vibrotactile Feedback & Skin Stimulation +1
- 60%
QButterfly: Lightweight Survey Extension for Online User-Interaction Studies for Non-Tech-Savvy Researchers
CHI '23· User Research Methods (Interviews, Surveys, Observation) +1
- 60%
Towards Fair and Equitable Incentives to Motivate Paid and Unpaid Crowd Contributions
CHI '25· Crowdsourcing Task Design & Quality Control +1
- 60%
Explainable Modeling of Annotations in Crowdsourcing
IUI '19· Explainable AI (XAI) +1
- 60%
Sprout: Crowd-Powered Task Design for Crowdsourcing
UIST '18· Crowdsourcing Task Design & Quality Control +1
Based on Jaccard similarity of research subtopics & professions (≥60%)