If in a Crowdsourced Data Annotation Pipeline, a GPT-4
Authors
Title of the Paper
If in a Crowdsourced Data Annotation Pipeline, a GPT-4
Paper Information
- Research Area: Artificial Intelligence, Crowdsourced Data Annotation, Natural Language Processing
- Keywords: Crowdsourcing, Data Annotation, GPT-4, Large Language Models (LLM), Data Cleaning, Label Aggregation, Human-AI Collaboration, Data Accuracy, Human-Computer Interaction, Data Quality
Research Background and Problems
-
Problems or Challenges:
- GPT-4 has been found to surpass crowdsourced workers, especially those from Amazon Mechanical Turk (MTurk), in terms of accuracy and efficiency in data annotation;
- Existing studies often overlook standard crowdsourcing practices and focus more on individual worker performance rather than the overall effectiveness of the data annotation pipeline;
- There is a lack of in-depth exploration on how to integrate GPT-4 and crowdsourced worker labels to improve overall accuracy.
-
Significance: High-quality data annotation is critical for training advanced AI models, but relying on humans for large-scale data annotation is costly and inefficient. Evaluating and optimizing the potential of combining GPT-4 with crowdsourced workers holds significant value.
-
Motivation and Related Work:
- Existing literature indicates that GPT-4 outperforms crowdsourced workers in limited scenarios, but often neglects label cleaning and aggregation in real-world contexts;
- This study focuses on a more comprehensive comparison between GPT-4 and standardized, ethically executed crowdsourced data annotation workflows, exploring the synergy between the two.
Solution
-
Methods or Solutions:
- Collected 127,080 labels from 415 MTurk workers and annotations from GPT-4, with tasks based on the CODA-19 labeling system to classify sentences and paragraphs in academic article abstracts;
- Designed two distinct worker interfaces (basic and advanced) to test the potential impact of interface design on annotation quality;
- Evaluated the quality of final labels using eight label aggregation algorithms (e.g., Majority Voting and Dawid-Skene);
- Combined GPT-4 annotations with crowdsourced worker labels to test the synergistic effects of their integration.
-
Innovations:
- Considered the complete crowdsourced data annotation pipeline, including data cleaning, quality control, and label aggregation techniques;
- Explored the complementarity of GPT-4 and crowdsourced worker labels, proposing that combining their strengths can improve overall accuracy.
-
Implementation Steps and Techniques:
- Data collection and grouping: Based on the CODA-19 labeling system, MTurk workers provided 20 labels per article;
- Interface differentiation experiment: Tested the efficiency and data quality of basic and advanced interfaces;
- Data cleaning and filtering: Applied strategies such as "All," "Exclude-By-Worker," and "Exclude-By-Batch";
- Executed label aggregation algorithms to compare the performance of workers and GPT-4.
Research Findings
-
Specific Results:
- Even under optimal conditions, the highest accuracy of the MTurk pipeline was 81.5%, slightly lower than GPT-4's 83.6%;
- When combining GPT-4 annotations with crowdsourced worker labels, two algorithms (One-Coin Dawid-Skene and MACE) achieved accuracy rates of 87.5% and 87.0%, respectively, surpassing the accuracy of using GPT-4 alone;
- The study revealed that crowdsourced workers performed better in certain specific categories (e.g., "Finding/Contribution"), complementing GPT-4's overall advantages.
-
Advantages Compared to Existing Solutions:
- Covered a more comprehensive process analysis, from data collection to subsequent data processing;
- Emphasized the potential of human-AI collaboration rather than relying solely on one method;
- Provided a broader methodological perspective by comparing various label cleaning and aggregation techniques.
-
Experimental or Evaluation Results:
- Under the "Exclude-By-Worker" strategy, the One-Coin Dawid-Skene and MACE algorithms significantly improved accuracy in scenarios combining GPT annotations;
- Experiments with mixed labels showed that crowdsourced labels optimized GPT's performance, particularly in subcategories where GPT was weaker.
-
Limitations and Future Directions:
- This study focused solely on the CODA-19 annotation task, and its generalizability to other tasks requires further validation;
- Did not test other LLM models (e.g., GPT-3.5) or open models (e.g., LLaMA);
- Future research directions: Further explore how to generate small-scale, high-quality expert-labeled data to optimize LLM annotation performance; design more robust user interfaces and interactions.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How does GPT-4 compare with crowd workers in data annotation performance?Category: Crowdsourcing Task Clarity and Quality ControlSimilar questionsarrow_forward
- How can combining GPT-4 and crowd worker labels improve data annotation accuracy?Category: Crowdsourcing Task Clarity and Quality ControlSimilar questionsarrow_forward
- How does interface design affect crowd workers' data annotation quality?Category: Crowdsourcing Task Clarity and Quality ControlSimilar questionsarrow_forward
Practical Problems
1- Large-scale high-quality data annotation is costly and inefficient.Category: Crowdsourcing Task Clarity and Quality ControlSimilar questionsarrow_forward
- 60%
Online Sequencing of Non-Decomposable Macrotasks in Expert Crowdsourcing
CHI '18· Crowdsourcing Task Design & Quality Control
- 60%
Directed Diversity: Leveraging Language Embedding Distances for Collective Creativity in Crowd Ideation
CHI '21· Generative AI (Text, Image, Music, Video) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)