The ``Colonial Impulse" of Natural Language Processing: An Audit of Bengali Sentiment Analysis Tools and Their Identity-based Biases
Authors
Title of the Paper
The “Colonial Impulse” of Natural Language Processing: An Audit of Bengali Sentiment Analysis Tools and Their Identity-based Biases
Paper Information
- Subject Area: Natural Language Processing (NLP), Algorithmic Bias, Postcolonial Studies, Human-Computer Interaction (HCI)
- Keywords: Algorithm Audit, Sentiment Analysis Tools, Bias, Identity, Coloniality
Research Background and Problem
-
Identified Issues or Challenges:
- NLP tools may reproduce colonial biases and values, yet there is a lack of research on the mechanisms through which these biases operate.
- Significant inequality in language resources and research: for example, English has far more resources than Bengali, despite the two languages having comparable numbers of speakers.
- Biases in non-English NLP systems and their impacts remain underexplored.
-
Significance:
- Sentiment analysis tools are widely used in practice, such as in content moderation and public opinion analysis, but identity biases within these tools may reinforce existing social injustices.
- Given the complex interplay of colonial history, religion, gender, and ethnicity in South Asian Bengali culture, this region requires a specific research framework.
-
Research Motivation:
- To investigate whether and how existing Bengali sentiment analysis tools exhibit biases based on gender, religion, and ethnicity.
- To understand how these biases relate to the demographic background of tool developers.
Solution
-
Proposed Approach:
- Conduct a systematic audit of Bengali sentiment analysis tools available on the Python Package Index (PyPI) and GitHub to examine identity-based biases.
- Utilize the existing Bengali Identity Bias Evaluation Dataset (BIBED), which contains sentences capable of explicitly or implicitly expressing identity.
-
Innovations:
- Combine postcolonial studies with algorithm auditing to analyze how sentiment analysis tools reflect colonial social structures.
- Introduce a dual-layer analysis method of implicit and explicit identity expressions to explore the tools' adaptability to complex social preferences.
- Quantify biases across multiple dimensions (e.g., gender, religion, and ethnicity) within the same tool.
-
Implementation Steps and Techniques:
- Tool Selection: Screened 13 sentiment analysis tools from PyPI and GitHub based on functionality and documentation completeness.
- Dataset Construction: Tested tool outputs using sentences from the BIBED dataset.
- Algorithm Audit: Input identical sentences (with variations in key identity terms) into different tools and statistically analyzed their sentiment scores.
- Statistical Analysis: Used Kruskal-Wallis and Mann-Whitney U tests to determine whether score differences were significant.
- Correlation Between Bias and Developer Background: Explored the relationship between published developer background information and the biases exhibited by the tools.
Research Findings
-
Specific Findings:
- Inconsistencies: Significant differences in sentiment scores for the same sentences across tools, indicating a lack of universality in sentiment measurement.
- Differences in Implicit vs. Explicit Expressions: Sentences with explicit identity mentions (e.g., direct references to “Bangladesh”) were often assigned more negative sentiments, while implicit expressions using regional dialects scored higher.
- Identity Bias:
- 38% of tools favored female identities, while 30% favored male identities.
- 30% of tools exhibited bias toward Hindu identities, while 38% favored Muslim identities.
- 77% of tools were biased toward Bangladeshi identities, with only 15% favoring Indian identities.
- Developer Background Correlation: While no direct relationship was observed between developer background and tool bias, the homogeneity of developers' ethnic and religious identities suggests potential limitations in development perspectives.
-
Advantages Over Existing Solutions:
- Conducted the first systematic algorithm audit of Bengali sentiment analysis tools.
- Explored differences between explicit and implicit identity expressions in sentiment analysis for the first time.
- Introduced the context of South Asian colonial history into the study of algorithmic bias, expanding the scope of bias research.
-
Limitations and Future Directions:
- Binary Identity Representation: This study only considered three binary classifications: male-female, Hindu-Muslim, and Bangladeshi-Indian, excluding broader groups (e.g., transgender individuals, other minority religions).
- Intersectional Bias Analysis: Did not explore mechanisms of bias in intersectional identities (e.g., Muslim women, transgender Indians).
- Future Work:
- Use qualitative methods (e.g., interviews) to delve deeper into how tools handle more nuanced identity categories.
- Investigate how different datasets contribute to the construction of tool biases.
- Expand to other low-resource languages and analyze sampling biases related to other social dimensions (e.g., caste, sexual orientation).
This study highlights the limitations of Bengali sentiment analysis tools and their potential societal impacts while offering a method to incorporate postcolonial perspectives into technology research.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- Do existing Bengali sentiment analysis tools exhibit bias based on gender, religion, and ethnicity?Category: Gender, Sexuality Bias, and Women/LGBTQ+ Experiences in AI, Technology, and Online PlatformsSimilar questionsarrow_forward
- How are biases in these tools related to developers' ethnic and religious backgrounds?Category: Gender, Sexuality Bias, and Women/LGBTQ+ Experiences in AI, Technology, and Online PlatformsSimilar questionsarrow_forward
- How do sentiment analysis tools handle differences between explicit and implicit identity expressions?Category: Gender, Sexuality Bias, and Women/LGBTQ+ Experiences in AI, Technology, and Online PlatformsSimilar questionsarrow_forward
Practical Problems
1- Bengali sentiment analysis tools may deepen existing social injustices.Category: Gender, Sexuality Bias, and Women/LGBTQ+ Experiences in AI, Technology, and Online PlatformsSimilar questionsarrow_forward
- 80%
Partiality and Misconception: Investigating Cultural Representativeness in Text-to-Image Models
CHI '24· AI Ethics, Fairness & Accountability +1
- 67%
Conceptualizing Algorithmic Stigmatization
CHI '23· AI Ethics, Fairness & Accountability +2
- 67%
When Stereotypes GTG: The Impact of Predictive Text Suggestions on Gender Bias in Human-AI Co-Writing
CHI '26· Human-LLM Collaboration +2
- 67%
Beyond Microsoft and Monsanto: Denaturing the Monoculture Metaphor in Computing
CHI '26· AI Ethics, Fairness & Accountability +2
- 60%
The Landscape and Gaps in Open Source Fairness Toolkits
CHI '21· AI Ethics, Fairness & Accountability +1
- 60%
Towards Fairness in Practice: A Practitioner-Oriented Rubric for Evaluating Fair ML Toolkits
CHI '21· AI Ethics, Fairness & Accountability +1
- 60%
Jury Learning: Integrating Dissenting Voices into Machine Learning Models
CHI '22· AI Ethics, Fairness & Accountability +1
- 60%
Capable but Amoral? Comparing AI and Human Expert Collaboration in Ethical Decision Making
CHI '22· AI Ethics, Fairness & Accountability +1
- 60%
Out of Context: Investigating the Bias and Fairness Concerns of "Artificial Intelligence as a Service"
CHI '23· AI Ethics, Fairness & Accountability +1
- 60%
“It is currently hodgepodge”: Examining AI/ML Practitioners’ Challenges during Co-production of Responsible AI Values
CHI '23· AI Ethics, Fairness & Accountability +1
Based on Jaccard similarity of research subtopics & professions (≥60%)