LabelAId: Just-in-time AI Interventions for Improving Human Labeling Quality and Domain Knowledge in Crowdsourcing Systems

Explainable AI (XAI)Crowdsourcing Task Design & Quality ControlUI/UX DesignersHCI ResearchersAmazon Mechanical Turk Workers

Document Title

LabelAId: Just-in-time AI Interventions for Improving Human Labeling Quality and Domain Knowledge in Crowdsourcing Systems

Document Information

  • Subject Area: Human-Computer Interaction, Machine Learning, Crowdsourcing Systems
  • Keywords: Crowdsourcing, Citizen Science, Quality Control, Machine Learning, Programmatic Weak Supervision (PWS), Urban Accessibility, Human-AI Collaboration

Research Background and Problem

  • What problems or challenges did the authors identify? Crowdsourcing systems have revolutionized the resolution of distributed problems, but data quality control remains a major challenge. Current methods (e.g., worker screening and vote filtering) focus on optimizing economic output, with limited impact on improving data quality and participant learning. Additionally, in citizen science tasks, where participants are often volunteers lacking domain expertise, data errors are even more pronounced.

  • Why is this problem important? High-quality data is critical for scientific research and AI model training. Simultaneously, enhancing participants' domain knowledge can encourage broader public engagement in scientific research.

  • Research Motivation and Related Work Existing quality control and learning feedback methods require additional expert support and are difficult to scale. LabelAId aims to address these issues through AI interventions, combining human-AI collaboration with real-time feedback to improve data labeling quality and participant learning experiences.

Solution

  • What methods or solutions did the authors propose?

    1. LabelAId is a system based on programmatic weak supervision (PWS) and machine learning pipelines, designed to dynamically infer labeling errors through human behavior and domain knowledge.
    2. A real-time system integration provides AI-based feedback when users make errors.
  • What is innovative about this solution?

    1. The PWS-based machine learning approach efficiently leverages unlabeled data to generate trainable data, reducing the need for manual intervention.
    2. The system offers real-time labeling assistance while enhancing user learning opportunities, making it more efficient compared to existing interaction feedback methods that require human support.
  • What are the implementation steps and key technologies used?

    1. PWS Label Generation:
      • Automated probabilistic labeling data is generated using domain knowledge and user behavior.
    2. Model Pretraining and Fine-tuning:
      • Machine learning models are pretrained using the generated automated labeling data.
      • Fine-tuning is performed on a small amount of manually validated data.
    3. Real-time Feedback and User Interface Design:
      • A decision-making interface with cognitive nudging features is designed, where users label first and then receive related feedback, such as typical errors and correct examples.

Research Outcomes

  • What specific outcomes were achieved?

    • LabelAId significantly improved labeling accuracy (e.g., a 19.2% increase in accuracy for Curb Ramp labels) without a notable decline in labeling speed.
    • Technical evaluations demonstrated that LabelAId's inference accuracy in low-data scenarios outperformed several mainstream machine learning models (e.g., XGBoost, MLP) by 36.7%.
    • In terms of user learning and confidence, LabelAId's feedback had a positive impact on participants' understanding of the task domain.
  • What advantages does it have compared to existing solutions?

    • Compared to manual verification or peer feedback, it reduces additional human labor costs.
    • It implements an efficient, scalable human-AI interactive feedback method that is easier to deploy in new crowdsourcing application domains.
  • What were the experimental or evaluation results?

    • LabelAId's positive impact on task performance, user learning, and confidence was widely reflected in technical evaluations and user studies. Examples include:
      • With 50 downstream labeled data points, misclassification inference accuracy exceeded leading ML models by 15%-36.7%.
      • It maintained generalizability in scenarios involving cities not included in training, achieving performance comparable to tasks in multiple cities.
  • Limitations and Future Directions

    • The system's inference of labeling errors still exhibits category-specific differences, such as lower performance in predicting errors for the "Obstacle" label.
    • No significant correlation was found between the frequency of user-AI interactions and improvements in labeling performance, indicating a need to further optimize the interaction interface and intervention frequency.
    • Suggestions for future work include:
      1. Enhancing the interpretability and transparency of model error feedback.
      2. Providing more flexible user control options, such as reducing or dynamically adjusting AI feedback frequency.
      3. Expanding to other crowdsourcing tasks requiring domain expertise, such as medical image annotation and wildlife classification.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147498/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642089
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
10 authors
sell
Subtopics
Explainable AI (XAI), Crowdsourcing Task Design & Quality Control
work
Professions
UI/UX Designers, HCI Researchers, Amazon Mechanical Turk Workers
article
Content Status
Full text indexed
hub
Related Papers
3 related papers