Designing Ground Truth and the Social Life of Labels

Computational Methods in HCIResearch Ethics & Open ScienceAI/ML Researchers & EngineersHCI Researchers

Title of the Paper

Designing Ground Truth and the Social Life of Labels

Paper Information

  • Research Domain: Artificial Intelligence and Human-Computer Interaction, focusing on data labeling processes and practices in machine learning.
  • Keywords: Human-Centered Data Science, Data Labeling, Label Design, Collaborative Computing, Machine Learning, Data Quality, Social Computing, Expertise, Work Practices, HCI

Research Background and Issues

  • Issues and Challenges:

    • Current research on "ground truth" data label design and annotation primarily focuses on crowdsourced workers, with limited attention to the specialized work practices of domain experts.
    • In machine learning models, data preparation processes account for a significant amount of time (up to 80%), with the annotation process being labor-intensive and complex, particularly in scenarios requiring high-quality labels.
    • The social and collaborative aspects of label design remain underexplored.
  • Significance:

    • High-quality labels are critical for improving machine learning model performance.
    • The label design process in machine learning workflows has profound implications but is often overlooked as it is "hidden" within foundational data.
  • Research Motivation and Related Work:

    • The authors aim to provide a detailed description of team work practices involving domain experts, exploring the social dynamics of label design and creation.
    • Previous studies have focused on crowdsourcing methods and automated labeling techniques, with limited exploration of human collaborative annotation in scenarios requiring high levels of expertise.

Solution

  • Methods and Solutions:

    • Employ qualitative research methods, including 15 in-depth interviews, summarizing descriptions of data science practitioners.
    • Identify three modes of ground truth label design:
      1. Principled Design: Based on predefined processes with detailed planning.
      2. Iterative Design: Improving label definitions through multiple trials and gradual adjustments.
      3. Improvisational Design: Making ad-hoc adjustments to address unforeseen challenges.
  • Innovations:

    • Integrates the label design process with Human-Centered Data Science theories, emphasizing the social nature and flexibility of labeling practices.
    • Provides discussions on label quality control and analyzes various strategies for improving annotation processes.
  • Implementation Steps:

    • Collect data through interviews, analyzing data science teams' annotation tools, experiences, collaboration strategies, and label quality management processes.
    • Apply grounded theory methods to systematically analyze interview data, ultimately developing a categorized explanatory framework.

Research Findings

  • Specific Findings:

    • Label design and annotation processes vary depending on the environment, resources, and personnel, but can generally be categorized into three main modes: principled, iterative, and improvisational.
    • Labels not only describe the world but also reflect social needs such as collaboration, time, and resource constraints.
    • Highlights the importance of team collaboration in managing label quality, such as resolving disagreements and improving label consistency.
    • Proposes several design recommendations for improving labeling tools and processes, including intelligent user interfaces, socialized tools, and label traceability mechanisms.
  • Advantages Compared to Existing Solutions:

    • Provides a systematic analysis of specific issues and solutions within domain expert teams, rather than focusing solely on crowdsourcing scenarios.
    • Introduces a comprehensive methodological framework in the field of data labeling, interpreting how team collaboration impacts label quality and effectiveness.
  • Experimental and Evaluation Results:

    • Summarizes the rich practical experiences of interviewees, showcasing recent challenges and solutions in the field of annotation and data preparation.
    • Finds that annotation quality is deeply influenced by annotators' roles, background knowledge, and collaboration strategies.
  • Limitations and Future Directions:

    • Limitations:
      • The study is primarily based on IBM teams in North America, lacking broader international data support.
      • The sample size is limited and focuses mainly on research-oriented projects; development or product-oriented teams may exhibit different behaviors.
      • Direct interviews with annotation operators (e.g., individual annotators) were not conducted, leading to certain perspective limitations.
    • Future Directions:
      • Further research into the specific contributions of different roles (e.g., annotators and project managers) in the annotation process.
      • Development of intelligent support tools, such as labeling tools that dynamically adjust task difficulty and socialized user interfaces to facilitate collaboration.
      • Expand discussions on label quality, exploring best practices for combining automated and human annotation.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/47890/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3411764.3445402
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2021
emoji_events
Award
No award tagged
group
Authors
11 authors
sell
Subtopics
Computational Methods in HCI, Research Ethics & Open Science
work
Professions
AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers