Increasing the Speed and Accuracy of Data Labeling Through an AI Assisted Interface
Authors
Title of the Paper
Increasing the Speed and Accuracy of Data Labeling Through an AI Assisted Interface
Paper Information
- Subject Area: Artificial Intelligence and Human-Computer Interaction; AI Collaboration in Data Labeling
- Keywords: AI Assistance, Data Labeling, Human-Computer Interaction, Machine Learning, Intelligent Interface Design
Research Background and Problem
-
Identified Problems or Challenges:
- Data labeling is a critical yet tedious step in supervised learning, requiring extensive human decision-making.
- Labeling tasks with large-scale label sets (ranging from tens to thousands of labels) are highly complex, often resulting in reduced accuracy and increased time consumption.
- Existing studies primarily focus on the effects of AI assistance in binary classification scenarios, with limited exploration of AI assistance in multi-label scenarios.
-
Significance of the Problem:
- As AI systems are increasingly utilized, the quality of labeled datasets directly impacts the performance of machine learning models. Improving the accuracy and efficiency of data labeling has become a critical need.
-
Research Motivation and Related Work:
- Human-AI collaboration and AI-assisted technologies have been proven to enhance the accuracy of binary decision-making tasks.
- Existing literature has discussed the impact of transparency, explainability, and trust calibration on AI-assisted decision-making, but there remains a significant gap in the design of AI assistance and user behavior studies for multi-label tasks.
Proposed Solution
-
Method or Solution:
The authors propose an AI-assisted data labeling system based on a semi-supervised learning algorithm to predict the potential label distribution of samples and provide labeling assistance through the following methods:- Offering label recommendations;
- Reducing the decision space for labelers, focusing only on the most likely labels.
-
Innovative Contributions:
- Addressing complex decision-making problems in multi-label scenarios;
- Proposing a labeling interface design that ranks and reduces the decision space based on predicted probabilities;
- Designing comprehensive experiments to evaluate the specific impact of AI assistance on labeling accuracy and speed.
-
Implementation Steps and Key Techniques:
- Using a semi-supervised learning algorithm based on Label Spreading to predict the label distribution of unlabeled samples.
- Displaying the top 5 most likely labels on the labeling interface, ranked by predicted probabilities.
- Providing a "view all labels" feature to allow user intervention.
- Designing experiments to evaluate the assistance effect by comparing the performance of baseline, weak AI, and strong AI groups.
Research Findings
-
Specific Outcomes:
- AI assistance significantly improved labelers' accuracy: accuracy increased by 6% compared to the baseline group.
- AI assistance significantly reduced the time required to complete labeling tasks, especially when the needed label was among the AI-recommended "top 5 labels."
- There was no significant difference in accuracy between weak AI and strong AI assistance: users were able to identify incorrect predictions and make independent decisions.
- The distribution of AI recommendation confidence (predicted probabilities) influenced user behavior: high-confidence incorrect predictions from weak AI reduced user trust but increased critical judgment.
-
Advantages Compared to Existing Solutions:
- Applicable to more complex multi-label labeling scenarios, achieving higher efficiency than traditional methods.
- Effectively combines label recommendation and decision space reduction techniques, enhancing user experience.
-
Experimental and Evaluation Results:
The average accuracy of users in the weak AI-assisted group was 0.79, in the strong AI-assisted group was 0.78, and in the baseline group was only 0.72. The labeling time in the assisted groups was significantly shorter than in the baseline group. -
Limitations and Future Directions:
- The current study only covers one labeling task and should be extended to different task types and difficulty scenarios.
- The experiments involved non-expert labelers; further exploration is needed for the effects of assistance in expert labeling contexts.
- The presentation of labels in the user interface (e.g., whether to fix spatial order) requires further optimization and research.
- The potential impact of different AI system confidence distributions on user behavior warrants deeper exploration.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can AI-assisted systems improve accuracy and efficiency in multi-label data annotation tasks?Category: Human-in-the-Loop Labeling and Example SelectionSimilar questionsarrow_forward
- How do the ranking and spatial display of AI-recommended labels affect users' annotation decisions?Category: Uncertainty in Active Learning and Machine TeachingSimilar questionsarrow_forward
- How does the distribution of AI system prediction confidence affect users' trust and judgment?Category: Uncertainty in Active Learning and Machine TeachingSimilar questionsarrow_forward
Practical Problems
1- Complex multi-label data annotation is time-consuming and error-prone, affecting model performance.Category: Uncertainty in Active Learning and Machine TeachingSimilar questionsarrow_forward
- 71%
"Are You Really Sure?'' Understanding the Effects of Human Self-Confidence Calibration in AI-Assisted Decision Making
CHI '24· Explainable AI (XAI) +1
- 71%
DynamicLabels: Supporting Informed Construction of Machine Learning Label Sets with Crowd Feedback
IUI '24· AI-Assisted Decision-Making & Automation +1
- 63%
Evaluating the Interpretability of Generative Models by Interactive Reconstruction
CHI '21· Explainable AI (XAI) +2
- 63%
PaTAT: Human-AI Collaborative Qualitative Coding with Explainable Interactive Rule Synthesis
CHI '23· Explainable AI (XAI) +2
- 63%
Faulty or Ready? Handling Failures in Deep-Learning Computer Vision Models until Deployment: A Study of Practices, Challenges, and Needs
CHI '23· Explainable AI (XAI) +2
- 63%
"Should I Rely on You or the AI?" Leaders' Trust and Perceptions in Mixed Human-AI Teams
CHI '26· Human-Robot Collaboration (HRC) +2
- 63%
Do People Appropriately Rely on AI-Advice? An Analytical Review of HCI Research on Human-AI Decision-Making
CHI '26· AI-Assisted Decision-Making & Automation +2
- 63%
Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving Tasks
CHI '26· Human-LLM Collaboration +2
- 63%
The Impact of Response Latency and Task Type on Human-LLM Interaction and Perception
CHI '26· Human-LLM Collaboration +2
- 63%
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
CHI '26· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)