Wikibench: Community-Driven Data Curation for AI Evaluation on Wikipedia

Community Collaboration & WikipediaCrowdsourcing Task Design & Quality ControlContent CuratorsHCI Researchers

Title of the Paper

Wikibench: Community-Driven Data Curation for AI Evaluation on Wikipedia

Bibliographic Information

  • Subject Area: Community-Driven Data Annotation and AI Evaluation
  • Keywords: Community-Driven AI, Data Annotation, AI Evaluation, Data Uncertainty, Wikipedia, Fairness, Human-AI Collaboration, Data Transparency, Label Consistency, Machine Learning Models

Research Background and Problem Statement

  • What problems or challenges did the authors identify?

    • Artificial Intelligence (AI) tools are increasingly deployed in community environments like Wikipedia. In these settings, AI evaluation datasets are often created by developers or external annotators, which may lead to misleading conclusions about AI performance.
    • Some existing datasets systematically fail to reflect community-specific norms and values, such as mislabeling non-offensive language as harmful content.
    • Current community methods for evaluating the applicability of AI tools are limited and lack the necessary infrastructure to support community-driven dataset curation.
  • Why is this problem important?

    • The quality of datasets directly impacts AI evaluation results and post-deployment performance. Misleading evaluations can degrade user experience or harm vulnerable communities.
    • Enhancing community consistency in data annotation helps better reflect societal complexity and cultural diversity, avoiding biases in tools.
  • Motivation and Related Work

    • Inspired by the community archiving tradition in library studies, which emphasizes the voices of non-elite and marginalized groups, the authors propose new methods for AI data annotation.
    • Existing AI evaluation techniques often address label disagreements through post-processing but lack mechanisms to facilitate discussion and negotiation among community members.

Solution

  • What methods or solutions did the authors propose?

    • The authors developed a new system—Wikibench—that enables Wikipedia community members to collaboratively annotate AI evaluation datasets, resolve label disputes, and address uncertainty.
    • The system integrates into Wikipedia's existing workflows through plugins, entity pages, and campaign pages.
  • What are the innovative aspects of this solution?

    • The system explores community-driven data curation strategies, empowering communities to decide which data points should be annotated and how they should be labeled.
    • It enables communities to collaboratively form a consensus primary label while preserving independent labels that represent individual perspectives.
    • The system supports community members in jointly defining label definitions and data inclusion standards, ensuring maximum transparency.
  • What are the implementation steps and key technologies used?

    • The plugin integrates with the edit review page, allowing seamless incorporation of labeling into routine activities.
    • Entity pages display all labels for a data point and support discussion, leveraging Wikipedia's "Talk Page" mechanism to resolve disagreements.
    • Campaign pages make the entire dataset public, highlighting label disagreements or data points requiring further attention, guiding discussions and additional annotations.
    • The design ensures data collaboration aligns with Wikipedia's community norms, such as the "Bold, Revert, Discuss (BRD)" cycle.

Research Outcomes

  • What specific outcomes were achieved?

    • Wikibench successfully collected 757 Wikipedia edit data points, creating a dataset that reflects community consensus, disagreements, and uncertainties.
    • Validation studies showed that primary labels generated by Wikibench exhibit greater community consistency compared to traditional tools like Wikilabels.
    • The dataset was used to evaluate the community consistency of two Wikipedia AI models (ORES and Revert-Risk), revealing that ORES aligns more closely with community perspectives.
  • What advantages does it have over existing solutions?

    • Compared to Wikilabels' linear annotation mechanism, Wikibench supports collaborative annotation and community-driven label definitions, better representing community consensus.
    • It captures label disagreements and uncertainties within the data, enabling more nuanced AI evaluation methods.
  • What were the experimental or evaluation results?

    • The primary labels generated by Wikibench demonstrated higher consistency with validation labels annotated through Wikipedia's standard negotiation process.
    • The ORES model performed better on the community-driven dataset compared to the Revert-Risk model, which was trained on historical edit data.
  • Limitations and Future Directions

    • The dataset does not fully represent the distribution of Wikipedia edits, which may limit its applicability for certain AI evaluation purposes.
    • Further exploration is needed to balance user participation levels with time efficiency.
    • Future research could explore extending Wikibench to other community environments, such as Reddit and Mastodon, and developing community visualization tools to support more complex AI evaluations.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/146778/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642278
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
8 authors
sell
Subtopics
Community Collaboration & Wikipedia, Crowdsourcing Task Design & Quality Control
work
Professions
Content Curators, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
1 related papers