Jury Learning: Integrating Dissenting Voices into Machine Learning Models

Best Paper
AI Ethics, Fairness & AccountabilityAlgorithmic Fairness & BiasAI/ML Researchers & EngineersHCI Researchers

Title of the Paper

Jury Learning: Integrating Dissenting Voices into Machine Learning Models

Paper Information

  • Subject Area: Human-Computer Interaction, Machine Learning, Fairness, and Social Computing
  • Keywords: Supervised Learning, AI Fairness, Annotation Disagreement, Jury Simulation, Diversity Modeling, Interactive Machine Learning, Custom Classifiers, Human-AI Collaboration, Classifier Interpretability

Research Background and Problem

  1. Problems and Challenges:

    • Current supervised learning primarily resolves annotation disagreements through majority voting, disregarding minority perspectives. This approach may fail to accurately capture the diversity of annotation disputes.
    • Machine learning models are often required to make judgments on societal issues, which frequently involve irreconcilable disagreements over the "correct label," such as toxicity detection in online comments, misinformation, and medical diagnoses.
    • Mainstream machine learning methods lack mechanisms to reflect minority opinions, and existing metrics often overlook who disagrees and why.
  2. Significance:

    • Training models solely based on majority perspectives in data may reinforce existing social inequalities and exacerbate biases.
    • A method is needed to enable machine learning systems to transparently reflect societal values as defined by relevant stakeholders.
  3. Research Motivation:

    • Propose a new mechanism to model annotators and allow decision-makers to explore annotation disagreements, helping researchers define customized classification rules to meet diverse needs.
    • Overcome the limitations of traditional methods by building transparent, dynamically adjustable machine learning tools for contentious tasks.

Solution

Method or Solution

  • Jury Learning:
    • Propose a supervised learning framework that models data annotators as jury members, allowing users to explicitly specify whose labels should guide model predictions by defining the jury composition (e.g., race, gender, political stance).
    • The model performs jury inference through sampling and outputs classification results based on weighted integration of jury opinions.

Innovations

  1. Dynamically Adjustable Jury:

    • Practitioners can dynamically adjust the jury composition to match task characteristics or cultural changes.
    • Enables exploration of counterfactual juries by altering jury composition to study decision changes.
  2. Annotator and Disagreement Modeling:

    • Leverages deep learning architectures (e.g., BERT and Deep & Cross Network) to model each annotator and predict individual labels for new inputs.
  3. Interactive Customization Interface:

    • Developed a visualization tool integrated with the Jury Learning interface, supporting interactive system interpretability by displaying annotator characteristics, sources of disagreement, and the "range of contention" in predictions.

Implementation Steps and Techniques

  1. Model Design:

    • Embed text into a deep recommendation system to form joint embeddings of annotators, inter-group relationships, and content representations.
    • Use a combined Deep & Cross Network to predict individual annotation behavior.
  2. Extended Features:

    • Support task-dependent jury structures (e.g., adaptively adjusting jury composition based on comment context).
    • Provide optimization tools to automatically identify the minimal jury adjustments needed to reverse prediction outcomes.

Research Outcomes

  1. Specific Results:

    • The Jury Learning model significantly improved individual annotation prediction performance, reducing the mean absolute error (MAE) from 0.90 to 0.61 compared to existing methods.
    • The model more accurately predicted individual and group distribution ranges, reducing prediction bias across groups.
  2. Experimental Results:

    • In user-supervised results, diverse jury configurations constructed by participants provided greater diversity compared to the default dataset's racially homogeneous composition (74% White annotators). Non-White annotators increased by 2.9 times, and non-binary gender annotators increased by 31.5 times.
    • After configuring the jury, approximately 14% of classification results were adjusted (e.g., from "toxic" to "non-toxic" or vice versa), particularly affecting cases with significant annotation disagreements.
  3. Practical Applications:

    • User testing (e.g., experiments with content moderators) validated that Jury Learning is user-friendly and can better serve various communities in complex social contexts.
  4. Limitations and Future Directions:

    • The current method is limited to known annotator group characteristics and requires further development of unsupervised methods to discover unknown groups or diverse opinions.
    • Calls for further research on the applicability of Jury Learning in fields such as medical diagnosis and collaborative design, as well as exploring fine-grained fairness optimization and robust learning mechanisms.
    • Current results are susceptible to annotator privacy risks, suggesting the integration of differential privacy methods to mitigate such risks.

In summary, this study proposes the Jury Learning model, enhancing the scalability of supervised learning in addressing diversity and disagreement issues. By redefining annotators as "jurors" and enabling customization flexibility, Jury Learning represents a significant advancement in fairness and user adjustability within machine learning.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/68851/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/abs/10.1145/3491102.3502004
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2022
emoji_events
Award
Best Paper
group
Authors
7 authors
sell
Subtopics
AI Ethics, Fairness & Accountability, Algorithmic Fairness & Bias
work
Professions
AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers