LabelAId: Just-in-time AI Interventions for Improving Human Labeling Quality and Domain Knowledge in Crowdsourcing Systems
Authors
Document Title
LabelAId: Just-in-time AI Interventions for Improving Human Labeling Quality and Domain Knowledge in Crowdsourcing Systems
Document Information
- Subject Area: Human-Computer Interaction, Machine Learning, Crowdsourcing Systems
- Keywords: Crowdsourcing, Citizen Science, Quality Control, Machine Learning, Programmatic Weak Supervision (PWS), Urban Accessibility, Human-AI Collaboration
Research Background and Problem
-
What problems or challenges did the authors identify? Crowdsourcing systems have revolutionized the resolution of distributed problems, but data quality control remains a major challenge. Current methods (e.g., worker screening and vote filtering) focus on optimizing economic output, with limited impact on improving data quality and participant learning. Additionally, in citizen science tasks, where participants are often volunteers lacking domain expertise, data errors are even more pronounced.
-
Why is this problem important? High-quality data is critical for scientific research and AI model training. Simultaneously, enhancing participants' domain knowledge can encourage broader public engagement in scientific research.
-
Research Motivation and Related Work Existing quality control and learning feedback methods require additional expert support and are difficult to scale. LabelAId aims to address these issues through AI interventions, combining human-AI collaboration with real-time feedback to improve data labeling quality and participant learning experiences.
Solution
-
What methods or solutions did the authors propose?
- LabelAId is a system based on programmatic weak supervision (PWS) and machine learning pipelines, designed to dynamically infer labeling errors through human behavior and domain knowledge.
- A real-time system integration provides AI-based feedback when users make errors.
-
What is innovative about this solution?
- The PWS-based machine learning approach efficiently leverages unlabeled data to generate trainable data, reducing the need for manual intervention.
- The system offers real-time labeling assistance while enhancing user learning opportunities, making it more efficient compared to existing interaction feedback methods that require human support.
-
What are the implementation steps and key technologies used?
- PWS Label Generation:
- Automated probabilistic labeling data is generated using domain knowledge and user behavior.
- Model Pretraining and Fine-tuning:
- Machine learning models are pretrained using the generated automated labeling data.
- Fine-tuning is performed on a small amount of manually validated data.
- Real-time Feedback and User Interface Design:
- A decision-making interface with cognitive nudging features is designed, where users label first and then receive related feedback, such as typical errors and correct examples.
- PWS Label Generation:
Research Outcomes
-
What specific outcomes were achieved?
- LabelAId significantly improved labeling accuracy (e.g., a 19.2% increase in accuracy for Curb Ramp labels) without a notable decline in labeling speed.
- Technical evaluations demonstrated that LabelAId's inference accuracy in low-data scenarios outperformed several mainstream machine learning models (e.g., XGBoost, MLP) by 36.7%.
- In terms of user learning and confidence, LabelAId's feedback had a positive impact on participants' understanding of the task domain.
-
What advantages does it have compared to existing solutions?
- Compared to manual verification or peer feedback, it reduces additional human labor costs.
- It implements an efficient, scalable human-AI interactive feedback method that is easier to deploy in new crowdsourcing application domains.
-
What were the experimental or evaluation results?
- LabelAId's positive impact on task performance, user learning, and confidence was widely reflected in technical evaluations and user studies. Examples include:
- With 50 downstream labeled data points, misclassification inference accuracy exceeded leading ML models by 15%-36.7%.
- It maintained generalizability in scenarios involving cities not included in training, achieving performance comparable to tasks in multiple cities.
- LabelAId's positive impact on task performance, user learning, and confidence was widely reflected in technical evaluations and user studies. Examples include:
-
Limitations and Future Directions
- The system's inference of labeling errors still exhibits category-specific differences, such as lower performance in predicting errors for the "Obstacle" label.
- No significant correlation was found between the frequency of user-AI interactions and improvements in labeling performance, indicating a need to further optimize the interaction interface and intervention frequency.
- Suggestions for future work include:
- Enhancing the interpretability and transparency of model error feedback.
- Providing more flexible user control options, such as reducing or dynamically adjusting AI feedback frequency.
- Expanding to other crowdsourcing tasks requiring domain expertise, such as medical image annotation and wildlife classification.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can real-time AI intervention improve data annotation quality in crowdsourcing systems?Category: Context-Aware Sampling and Low-Disruption NotificationsSimilar questionsarrow_forward
- Can AI feedback simultaneously improve participants' domain knowledge and task confidence?Category: Context-Aware Sampling and Low-Disruption NotificationsSimilar questionsarrow_forward
- How can probabilistic weak supervision (PWS) machine learning methods improve error inference accuracy in low-data scenarios?Category: Context-Aware Sampling and Low-Disruption NotificationsSimilar questionsarrow_forward
Practical Problems
1- Crowdsourcing participants often produce annotation errors due to lack of domain knowledge, affecting data quality.Category: Context-Aware Sampling and Low-Disruption NotificationsSimilar questionsarrow_forward
- 80%
Explainable Modeling of Annotations in Crowdsourcing
IUI '19· Explainable AI (XAI) +1
- 67%
From Text to Pixels: Enhancing User Understanding through Text-to-Image Model Explanations
IUI '24· Generative AI (Text, Image, Music, Video) +1
- 60%
Online Sequencing of Non-Decomposable Macrotasks in Expert Crowdsourcing
CHI '18· Crowdsourcing Task Design & Quality Control
Based on Jaccard similarity of research subtopics & professions (≥60%)