A hunt for the Snark: Annotator Diversity in Data Practices

Honorable Mention
Explainable AI (XAI)Algorithmic Fairness & BiasAI/ML Researchers & EngineersAmazon Mechanical Turk Workers

Title of the Paper

A hunt for the Snark: Annotator Diversity in Data Practices

Paper Information

  • Subject Area: Data annotation, diversity, and production practices in AI/ML datasets
  • Keywords: Data annotation, data work, machine learning, ML datasets, diversity, annotator diversity, data production

Research Background and Problem Statement

  • What issues or challenges did the authors identify?

    • Data diversity is recognized as a core element of responsible AI/ML, yet there is limited research on the diversity of annotators involved in dataset creation and its impact.
    • The subjectivity and diversity of annotators are often overlooked in AI/ML practices, with annotator diversity rarely prioritized in dataset production.
    • Existing practices are constrained by operational barriers, such as lack of transparency in annotator recruitment processes and difficulties in integrating annotator diversity.
  • Why is this issue important?

    • Data annotation is a critical step in building machine learning models, and annotators' subjectivity can affect label quality, potentially leading to unfair or inaccurate models.
    • Insufficient diversity may exacerbate biases or inequalities in AI/ML systems.
  • Research Motivation and Related Work

    • Extend the understanding of the impact of annotator diversity by integrating discussions on diversity in AI system design and studies on human annotator subjectivity.
    • Investigate the operational barriers in annotation practices and the underlying logic to explore alternative approaches.

Proposed Solutions

  • What methods or solutions did the authors propose?

    • Conducted 16 semi-structured interviews and collected 44 survey responses to study AI/ML practitioners' understanding and operationalization of annotator diversity.
    • Used the social theory of "regimes of existence" as an analytical framework to reflect on the neglect of annotator diversity in existing data practices.
    • Proposed rethinking "accurate data" (Ground Truth), bias, and diversity.
  • What is innovative about this solution?

    • Focused not only on increasing annotator diversity but also suggested rethinking data practices from an epistemic orientation, challenging the current logic centered on neutrality and objectivity.
    • Incorporated intersectionality theory to provide a deeper discussion on justice-oriented diversity.
  • Implementation Steps and Key Techniques

    1. Research Methods: Mixed methods approach, including surveys to gather perspectives and experiences on diversity, and semi-structured interviews to explore challenges in practice.
    2. Data Analysis: Quantitative analysis of survey results; qualitative analysis of interview content through iterative coding and thematic synthesis.
    3. Theoretical Framework: Utilized representationalist thinking, regimes of existence, and intersectionality theory as the foundation for critique and analysis.

Research Findings

  • What specific findings were achieved?

    1. Described the workflow of annotation work, where annotators are treated as instrumental "observers," with their personal perspectives and backgrounds removed in the name of "neutrality."
    2. Summarized three approaches to annotator diversity in AI/ML practices: ignoring diversity, pursuing objective standards, and attempting "neutral representation" through representativeness.
    3. Identified key barriers to annotator diversity, including:
      • Lack of background data on annotators.
      • Operational separation between annotation platforms and AI developers, leading to communication gaps.
      • The prioritization of model performance over diversity in machine learning development.
    4. Revealed through phenomenological analysis that current AI data practices are dominated by a "representationalist" logic, which diminishes the importance of diversity.
  • How does it compare to existing solutions?

    • Theoretically critiques and surpasses the singular pursuit of "objectivity" in current practices.
    • Provides justice-oriented and inclusive practical recommendations.
  • What are the experimental or evaluation results?

    • 75% of survey respondents believed annotator diversity significantly impacts dataset quality, yet it is rarely considered in actual recruitment or production processes.
    • Surveys and interviews exposed a disconnect between task design and considerations of diversity.
  • Limitations and Future Directions

    • The study sample may be biased toward practitioners already concerned with diversity; future research should expand the sample to include broader practices.
    • Future work could explore bridging communication gaps between annotation platforms and AI developers or designing tools to support diversity-oriented annotation practices.
    • Calls for more scholars to examine and test challenging issues within intersectionality, such as dynamic group formations and power structures.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/95795/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3580645
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
Honorable Mention
group
Authors
3 authors
sell
Subtopics
Explainable AI (XAI), Algorithmic Fairness & Bias
work
Professions
AI/ML Researchers & Engineers, Amazon Mechanical Turk Workers
article
Content Status
Full text indexed
hub
Related Papers
5 related papers