Explainable Modeling of Annotations in Crowdsourcing
Authors
Aggregation models for improving the quality of annotations collected via crowdsourcing have been widely studied, but far less has been done to explain why annotators make the mistakes that they do. To this end, we propose a joint aggregation and worker clustering model that detects patterns underlying crowd worker labels to characterize varieties of labeling errors. We evaluate our approach on a Named Entity Recognition dataset labeled by Mechanical Turk workers in both a retrospective experiment and a small human study. The former shows that our joint model improves the quality of clusters vs. aggregation followed by clustering. Results of the latter suggest that clusters aid human sense-making in interpreting worker labels and predicting worker mistakes. By enabling better explanation of annotator mistakes, our model creates a new opportunity to help Requesters improve task instructions and to help crowd annotators learn from their mistakes. Source code, data, and supplementary material is shared online.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 80%
LabelAId: Just-in-time AI Interventions for Improving Human Labeling Quality and Domain Knowledge in Crowdsourcing Systems
CHI '24· Explainable AI (XAI) +1
- 75%
Online Sequencing of Non-Decomposable Macrotasks in Expert Crowdsourcing
CHI '18· Crowdsourcing Task Design & Quality Control
- 67%
The Design and Development of a Game to Study Backdoor Poisoning Attacks: The Backdoor Game
IUI '21· Explainable AI (XAI) +2
- 60%
Crowdlicit: A System for Conducting Distributed End-User Elicitation and Identification Studies
CHI '19· Crowdsourcing Task Design & Quality Control +1
- 60%
Crowdsourced Detection of Emotionally Manipulative Language
CHI '20· AI Ethics, Fairness & Accountability +1
- 60%
Interface Design for Crowdsourcing Hierarchical Multi-Label Text Annotations
CHI '23· Crowdsourcing Task Design & Quality Control +1
- 60%
Can we crowdsource Tacton similarity perception and metaphor ratings?
CHI '23· Vibrotactile Feedback & Skin Stimulation +1
- 60%
Towards Fair and Equitable Incentives to Motivate Paid and Unpaid Crowd Contributions
CHI '25· Crowdsourcing Task Design & Quality Control +1
- 60%
Sprout: Crowd-Powered Task Design for Crowdsourcing
UIST '18· Crowdsourcing Task Design & Quality Control +1
Based on Jaccard similarity of research subtopics & professions (≥60%)