Wikibench: Community-Driven Data Curation for AI Evaluation on Wikipedia
Authors
Community Collaboration & WikipediaCrowdsourcing Task Design & Quality ControlContent CuratorsHCI Researchers
Title of the Paper
Wikibench: Community-Driven Data Curation for AI Evaluation on Wikipedia
Bibliographic Information
- Subject Area: Community-Driven Data Annotation and AI Evaluation
- Keywords: Community-Driven AI, Data Annotation, AI Evaluation, Data Uncertainty, Wikipedia, Fairness, Human-AI Collaboration, Data Transparency, Label Consistency, Machine Learning Models
Research Background and Problem Statement
-
What problems or challenges did the authors identify?
- Artificial Intelligence (AI) tools are increasingly deployed in community environments like Wikipedia. In these settings, AI evaluation datasets are often created by developers or external annotators, which may lead to misleading conclusions about AI performance.
- Some existing datasets systematically fail to reflect community-specific norms and values, such as mislabeling non-offensive language as harmful content.
- Current community methods for evaluating the applicability of AI tools are limited and lack the necessary infrastructure to support community-driven dataset curation.
-
Why is this problem important?
- The quality of datasets directly impacts AI evaluation results and post-deployment performance. Misleading evaluations can degrade user experience or harm vulnerable communities.
- Enhancing community consistency in data annotation helps better reflect societal complexity and cultural diversity, avoiding biases in tools.
-
Motivation and Related Work
- Inspired by the community archiving tradition in library studies, which emphasizes the voices of non-elite and marginalized groups, the authors propose new methods for AI data annotation.
- Existing AI evaluation techniques often address label disagreements through post-processing but lack mechanisms to facilitate discussion and negotiation among community members.
Solution
-
What methods or solutions did the authors propose?
- The authors developed a new system—Wikibench—that enables Wikipedia community members to collaboratively annotate AI evaluation datasets, resolve label disputes, and address uncertainty.
- The system integrates into Wikipedia's existing workflows through plugins, entity pages, and campaign pages.
-
What are the innovative aspects of this solution?
- The system explores community-driven data curation strategies, empowering communities to decide which data points should be annotated and how they should be labeled.
- It enables communities to collaboratively form a consensus primary label while preserving independent labels that represent individual perspectives.
- The system supports community members in jointly defining label definitions and data inclusion standards, ensuring maximum transparency.
-
What are the implementation steps and key technologies used?
- The plugin integrates with the edit review page, allowing seamless incorporation of labeling into routine activities.
- Entity pages display all labels for a data point and support discussion, leveraging Wikipedia's "Talk Page" mechanism to resolve disagreements.
- Campaign pages make the entire dataset public, highlighting label disagreements or data points requiring further attention, guiding discussions and additional annotations.
- The design ensures data collaboration aligns with Wikipedia's community norms, such as the "Bold, Revert, Discuss (BRD)" cycle.
Research Outcomes
-
What specific outcomes were achieved?
- Wikibench successfully collected 757 Wikipedia edit data points, creating a dataset that reflects community consensus, disagreements, and uncertainties.
- Validation studies showed that primary labels generated by Wikibench exhibit greater community consistency compared to traditional tools like Wikilabels.
- The dataset was used to evaluate the community consistency of two Wikipedia AI models (ORES and Revert-Risk), revealing that ORES aligns more closely with community perspectives.
-
What advantages does it have over existing solutions?
- Compared to Wikilabels' linear annotation mechanism, Wikibench supports collaborative annotation and community-driven label definitions, better representing community consensus.
- It captures label disagreements and uncertainties within the data, enabling more nuanced AI evaluation methods.
-
What were the experimental or evaluation results?
- The primary labels generated by Wikibench demonstrated higher consistency with validation labels annotated through Wikipedia's standard negotiation process.
- The ORES model performed better on the community-driven dataset compared to the Revert-Risk model, which was trained on historical edit data.
-
Limitations and Future Directions
- The dataset does not fully represent the distribution of Wikipedia edits, which may limit its applicability for certain AI evaluation purposes.
- Further exploration is needed to balance user participation levels with time efficiency.
- Future research could explore extending Wikibench to other community environments, such as Reddit and Mastodon, and developing community visualization tools to support more complex AI evaluations.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How can community-led data annotation be implemented in community environments (such as Wikipedia) to support AI evaluation?Category: Human-in-the-Loop Labeling and Example SelectionSimilar questionsarrow_forward
- Can community-driven dataset curation improve label consistency and reduce data bias?Category: Human-in-the-Loop Labeling and Example SelectionSimilar questionsarrow_forward
- What tools and mechanisms can support community members in jointly defining labeling standards, resolving annotation disputes, and improving data transparency?Category: Human-in-the-Loop Labeling and Example SelectionSimilar questionsarrow_forward
lightbulb
Practical Problems
1- AI evaluation datasets do not reflect community values, producing bias that affects user experience.Category: Human-in-the-Loop Labeling and Example SelectionSimilar questionsarrow_forward
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642278
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
8 authors
sell
Subtopics
Community Collaboration & Wikipedia, Crowdsourcing Task Design & Quality Control
work
Professions
Content Curators, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
1 related papers