Investigating the Capabilities and Limitations of Machine Learning for Identifying Bias in English Language Data with Information and Heritage Professionals

Honorable Mention
AI Ethics, Fairness & AccountabilityAlgorithmic Fairness & BiasMuseum Curators & Archivists

Research Background and Issues

  • What problems or challenges did the authors identify?
    Despite significant advancements in machine learning (ML), these systems still struggle to fully address social biases in language data. Specifically, when dealing with cultural heritage related to galleries, libraries, archives, and museums (GLAM), descriptive language often contains notable gender biases, which can influence researchers' and the public's understanding of history.

  • Why is this issue important?
    Gender bias in language can exacerbate social injustices and further harm marginalized groups. If machine learning fails to understand and manage these biases, its widespread application could amplify existing unequal power dynamics.

  • Research motivation and related work
    Existing ML methods primarily focus on "removing" biases from data or models. However, this approach overlooks the social and historical contexts of bias. Literature also indicates that many biases are inherently present in cultural heritage records. The authors aim to explore how to manage language biases in such cases through interdisciplinary methods, rather than attempting to eliminate them entirely.


Solution

  • What methods or solutions did the authors propose?
    The authors developed classification models using machine learning to identify gender bias in English language data. These models are designed to classify descriptive metadata from GLAM archives, tagging potentially biased language.

  • What is innovative about this solution?
    The primary innovation lies in redefining "bias." Instead of attempting to eliminate bias entirely, the authors propose using ML to reveal bias in language, making bias management a key objective. Additionally, the study employs a hybrid approach combining ML and human-computer interaction (HCI), emphasizing the critical role of humans in the ML workflow to ensure biases are correctly interpreted.

  • What are the implementation steps and key technologies used?

    • Data collection and preprocessing: Text data for training the models was collected from metadata in Scottish archives, annotated into 10 gender-related bias categories, including gendered pronouns, occupations, stereotypes, etc.
    • Model training: Three classifiers were developed using traditional supervised learning methods: a word-level classifier (LC), a sequence classifier (PNOC), and a document classifier (OSC), each targeting different types of linguistic phenomena.
    • Model cascading: Predictions from different classifiers were used as input features for other models, optimizing the final classifier's performance through cascading.
    • Human-computer interaction evaluation: Workshops were conducted to gather feedback from information and heritage professionals, deeply evaluating the models' ability to detect language bias and their effectiveness in real-world applications.

Research Outcomes

  • What specific results were achieved?

    • The ML models were able to classify gender-related language in GLAM metadata. The document classifier in Cascade 2 performed best in the "Stereotype" category (F1=0.841).
    • This study was the first to explore how gender bias manifests in GLAM metadata, demonstrating that the contextual and complex nature of bias exceeds the capabilities of traditional ML approaches.
  • What advantages does it have over existing solutions?

    • Unlike methods that solely focus on mathematically eliminating bias, this approach integrates sociological and critical theories, employing human oversight to manage bias.
    • It explores management strategies for unavoidable biases, broadening the discussion of bias and fairness in ML research.
  • What were the experimental or evaluation results?

    • Workshop results: Information and heritage professionals found that bias is contextual, dynamic, and unavoidable, which contrasts significantly with the technical definitions of bias in traditional ML research.
    • Model performance: While the models outperformed inter-annotator agreement among humans, certain categories (e.g., "Omission") performed poorly, indicating the need for further model design improvements.
  • Limitations and future directions

    • Limitations: Certain types of bias, especially implicit bias, are challenging to identify; model performance is limited in specific categories. There is insufficient data to support records of non-binary genders.
    • Future directions:
      • Expand datasets to better capture non-binary genders and complex bias patterns.
      • Conduct more experiments using deep learning while validating bias transfer with cost-effective traditional methods during the design phase.
      • Explore ways to reduce emotional stress during ML training and application, enhancing participants' confidence and feedback quality.

Overall, this study not only investigates how language bias manifests in the cultural heritage domain but also provides an innovative framework for how machine learning can address bias. It emphasizes the potential for ML to assist rather than fully automate bias management. This approach is not only applicable to the GLAM sector but can also be extended to other contexts with complex applications.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/189589/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713217
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
Honorable Mention
group
Authors
4 authors
sell
Subtopics
AI Ethics, Fairness & Accountability, Algorithmic Fairness & Bias
work
Professions
Museum Curators & Archivists
article
Content Status
Full text indexed
hub
Related Papers
2 related papers