If in a Crowdsourced Data Annotation Pipeline, a GPT-4

Generative AI (Text, Image, Music, Video)Crowdsourcing Task Design & Quality ControlData Scientists & AnalystsHCI ResearchersAmazon Mechanical Turk Workers

Title of the Paper

If in a Crowdsourced Data Annotation Pipeline, a GPT-4

Paper Information

  • Research Area: Artificial Intelligence, Crowdsourced Data Annotation, Natural Language Processing
  • Keywords: Crowdsourcing, Data Annotation, GPT-4, Large Language Models (LLM), Data Cleaning, Label Aggregation, Human-AI Collaboration, Data Accuracy, Human-Computer Interaction, Data Quality

Research Background and Problems

  • Problems or Challenges:

    1. GPT-4 has been found to surpass crowdsourced workers, especially those from Amazon Mechanical Turk (MTurk), in terms of accuracy and efficiency in data annotation;
    2. Existing studies often overlook standard crowdsourcing practices and focus more on individual worker performance rather than the overall effectiveness of the data annotation pipeline;
    3. There is a lack of in-depth exploration on how to integrate GPT-4 and crowdsourced worker labels to improve overall accuracy.
  • Significance: High-quality data annotation is critical for training advanced AI models, but relying on humans for large-scale data annotation is costly and inefficient. Evaluating and optimizing the potential of combining GPT-4 with crowdsourced workers holds significant value.

  • Motivation and Related Work:

    1. Existing literature indicates that GPT-4 outperforms crowdsourced workers in limited scenarios, but often neglects label cleaning and aggregation in real-world contexts;
    2. This study focuses on a more comprehensive comparison between GPT-4 and standardized, ethically executed crowdsourced data annotation workflows, exploring the synergy between the two.

Solution

  • Methods or Solutions:

    1. Collected 127,080 labels from 415 MTurk workers and annotations from GPT-4, with tasks based on the CODA-19 labeling system to classify sentences and paragraphs in academic article abstracts;
    2. Designed two distinct worker interfaces (basic and advanced) to test the potential impact of interface design on annotation quality;
    3. Evaluated the quality of final labels using eight label aggregation algorithms (e.g., Majority Voting and Dawid-Skene);
    4. Combined GPT-4 annotations with crowdsourced worker labels to test the synergistic effects of their integration.
  • Innovations:

    1. Considered the complete crowdsourced data annotation pipeline, including data cleaning, quality control, and label aggregation techniques;
    2. Explored the complementarity of GPT-4 and crowdsourced worker labels, proposing that combining their strengths can improve overall accuracy.
  • Implementation Steps and Techniques:

    1. Data collection and grouping: Based on the CODA-19 labeling system, MTurk workers provided 20 labels per article;
    2. Interface differentiation experiment: Tested the efficiency and data quality of basic and advanced interfaces;
    3. Data cleaning and filtering: Applied strategies such as "All," "Exclude-By-Worker," and "Exclude-By-Batch";
    4. Executed label aggregation algorithms to compare the performance of workers and GPT-4.

Research Findings

  • Specific Results:

    1. Even under optimal conditions, the highest accuracy of the MTurk pipeline was 81.5%, slightly lower than GPT-4's 83.6%;
    2. When combining GPT-4 annotations with crowdsourced worker labels, two algorithms (One-Coin Dawid-Skene and MACE) achieved accuracy rates of 87.5% and 87.0%, respectively, surpassing the accuracy of using GPT-4 alone;
    3. The study revealed that crowdsourced workers performed better in certain specific categories (e.g., "Finding/Contribution"), complementing GPT-4's overall advantages.
  • Advantages Compared to Existing Solutions:

    1. Covered a more comprehensive process analysis, from data collection to subsequent data processing;
    2. Emphasized the potential of human-AI collaboration rather than relying solely on one method;
    3. Provided a broader methodological perspective by comparing various label cleaning and aggregation techniques.
  • Experimental or Evaluation Results:

    1. Under the "Exclude-By-Worker" strategy, the One-Coin Dawid-Skene and MACE algorithms significantly improved accuracy in scenarios combining GPT annotations;
    2. Experiments with mixed labels showed that crowdsourced labels optimized GPT's performance, particularly in subcategories where GPT was weaker.
  • Limitations and Future Directions:

    1. This study focused solely on the CODA-19 annotation task, and its generalizability to other tasks requires further validation;
    2. Did not test other LLM models (e.g., GPT-3.5) or open models (e.g., LLaMA);
    3. Future research directions: Further explore how to generate small-scale, high-quality expert-labeled data to optimize LLM annotation performance; design more robust user interfaces and interactions.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/148078/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642834
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Crowdsourcing Task Design & Quality Control
work
Professions
Data Scientists & Analysts, HCI Researchers, Amazon Mechanical Turk Workers
article
Content Status
Full text indexed
hub
Related Papers
2 related papers