A hunt for the Snark: Annotator Diversity in Data Practices
Honorable MentionExplainable AI (XAI)Algorithmic Fairness & BiasAI/ML Researchers & EngineersAmazon Mechanical Turk Workers
Title of the Paper
A hunt for the Snark: Annotator Diversity in Data Practices
Paper Information
- Subject Area: Data annotation, diversity, and production practices in AI/ML datasets
- Keywords: Data annotation, data work, machine learning, ML datasets, diversity, annotator diversity, data production
Research Background and Problem Statement
-
What issues or challenges did the authors identify?
- Data diversity is recognized as a core element of responsible AI/ML, yet there is limited research on the diversity of annotators involved in dataset creation and its impact.
- The subjectivity and diversity of annotators are often overlooked in AI/ML practices, with annotator diversity rarely prioritized in dataset production.
- Existing practices are constrained by operational barriers, such as lack of transparency in annotator recruitment processes and difficulties in integrating annotator diversity.
-
Why is this issue important?
- Data annotation is a critical step in building machine learning models, and annotators' subjectivity can affect label quality, potentially leading to unfair or inaccurate models.
- Insufficient diversity may exacerbate biases or inequalities in AI/ML systems.
-
Research Motivation and Related Work
- Extend the understanding of the impact of annotator diversity by integrating discussions on diversity in AI system design and studies on human annotator subjectivity.
- Investigate the operational barriers in annotation practices and the underlying logic to explore alternative approaches.
Proposed Solutions
-
What methods or solutions did the authors propose?
- Conducted 16 semi-structured interviews and collected 44 survey responses to study AI/ML practitioners' understanding and operationalization of annotator diversity.
- Used the social theory of "regimes of existence" as an analytical framework to reflect on the neglect of annotator diversity in existing data practices.
- Proposed rethinking "accurate data" (Ground Truth), bias, and diversity.
-
What is innovative about this solution?
- Focused not only on increasing annotator diversity but also suggested rethinking data practices from an epistemic orientation, challenging the current logic centered on neutrality and objectivity.
- Incorporated intersectionality theory to provide a deeper discussion on justice-oriented diversity.
-
Implementation Steps and Key Techniques
- Research Methods: Mixed methods approach, including surveys to gather perspectives and experiences on diversity, and semi-structured interviews to explore challenges in practice.
- Data Analysis: Quantitative analysis of survey results; qualitative analysis of interview content through iterative coding and thematic synthesis.
- Theoretical Framework: Utilized representationalist thinking, regimes of existence, and intersectionality theory as the foundation for critique and analysis.
Research Findings
-
What specific findings were achieved?
- Described the workflow of annotation work, where annotators are treated as instrumental "observers," with their personal perspectives and backgrounds removed in the name of "neutrality."
- Summarized three approaches to annotator diversity in AI/ML practices: ignoring diversity, pursuing objective standards, and attempting "neutral representation" through representativeness.
- Identified key barriers to annotator diversity, including:
- Lack of background data on annotators.
- Operational separation between annotation platforms and AI developers, leading to communication gaps.
- The prioritization of model performance over diversity in machine learning development.
- Revealed through phenomenological analysis that current AI data practices are dominated by a "representationalist" logic, which diminishes the importance of diversity.
-
How does it compare to existing solutions?
- Theoretically critiques and surpasses the singular pursuit of "objectivity" in current practices.
- Provides justice-oriented and inclusive practical recommendations.
-
What are the experimental or evaluation results?
- 75% of survey respondents believed annotator diversity significantly impacts dataset quality, yet it is rarely considered in actual recruitment or production processes.
- Surveys and interviews exposed a disconnect between task design and considerations of diversity.
-
Limitations and Future Directions
- The study sample may be biased toward practitioners already concerned with diversity; future research should expand the sample to include broader practices.
- Future work could explore bridging communication gaps between annotation platforms and AI developers or designing tools to support diversity-oriented annotation practices.
- Calls for more scholars to examine and test challenging issues within intersectionality, such as dynamic group formations and power structures.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How does annotator diversity in data annotation affect dataset quality and fairness of machine learning models?Category: Machine Learning Fairness and Data Development PracticesSimilar questionsarrow_forward
- What in current AI/ML data practices hinders inclusion of annotator diversity?Category: Machine Learning Fairness and Data Development PracticesSimilar questionsarrow_forward
- How can data accuracy and diversity be redefined from a justice-oriented perspective?Category: Machine Learning Fairness and Data Development PracticesSimilar questionsarrow_forward
lightbulb
Practical Problems
1- AI dataset annotation often neglects annotator diversity, leading to model bias.Category: Machine Learning Fairness and Data Development PracticesSimilar questionsarrow_forward
- 75%
Designing Interactive Explainable AI Tools for Algorithmic Literacy and Transparency
DIS '24· Explainable AI (XAI) +1
- 67%
Explaining Models: An Empirical Study of How Explanations Impact Fairness Judgment
IUI '19· Explainable AI (XAI) +2
- 60%
"Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AI
CHI '21· Explainable AI (XAI) +1
- 60%
Fairness Evaluation in Text Classification: Machine Learning Practitioner Perspectives of Individual and Group Fairness
CHI '23· Explainable AI (XAI) +1
- 60%
Perceptions of the Fairness Impacts of Multiplicity in Machine Learning
CHI '25· Explainable AI (XAI) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3580645
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
Honorable Mention
group
Authors
3 authors
sell
Subtopics
Explainable AI (XAI), Algorithmic Fairness & Bias
work
Professions
AI/ML Researchers & Engineers, Amazon Mechanical Turk Workers
article
Content Status
Full text indexed
hub
Related Papers
5 related papers