Model Positionality and Computational Reflexivity: Promoting Reflexivity in Data Science

Honorable Mention
Explainable AI (XAI)AI Ethics, Fairness & AccountabilityPrivacy by Design & User ControlAI/ML Researchers & EngineersHCI ResearchersStatisticians & Data Scientists

Title of the Paper

Model Positionality and Computational Reflexivity: Promoting Reflexivity in Data Science

Bibliographic Information

  • Subject Areas: Human-Centered Data Science, Machine Learning Ethics, Annotated Data Processing
  • Keywords: Computational Reflexivity, Model Positionality, Data Science, Human-Centered Machine Learning, Annotator Fingerprinting, Position Mining, Critical Data Studies, Online Communities
  • Authors and Emails: Scott Allen Cambo (scottallencambo@gmail.com), Darren Gergle (dgergle@northwestern.edu)
  • Publication Venue and Date: CHI Conference on Human Factors in Computing Systems (CHI ’22), April 29-May 5, 2022
  • DOI: https://doi.org/10.1145/3491102.3501998

Research Background and Problem

  • Problems or Challenges Identified by the Authors:

    • Although data science and machine learning methods aim for objective analysis, subjective decisions during their development and deployment often go unacknowledged, leading to potential biases.
    • Models used on social platforms, such as algorithms for detecting online harassment, face practical limitations due to subjective human interpretations of concepts like "toxicity."
    • Large-scale machine learning workflows lack sufficient transparency regarding the sociocultural backgrounds of data scientists, annotators, and models.
  • Why This Problem is Important:

    • Data science is widely applied in critical areas affecting human lives (e.g., criminal justice, vaccine distribution), where potential biases and erroneous decisions can have profound implications for social justice.
    • A lack of deep reflection on the model development process can result in unfair treatment of minority groups and even systemic oppression.
  • Research Motivation and Related Work:

    • Building on the trend of human-centered data science, the authors aim to promote transparency and reflexive analysis in decision-making processes.
    • Drawing inspiration from qualitative research methods in social sciences, such as positionality and reflexivity, the authors extend these concepts to the context of data science.

Solution

  • Proposed Solution by the Authors:

    • Introduced two core concepts: "Model Positionality" and "Computational Reflexivity."
    • Developed new techniques: Annotator Fingerprinting and Position Mining, to help understand the sociocultural backgrounds of models, annotators, and data scientists.
  • Innovative Aspects of the Solution:

    • Model Positionality: Reveals the social and cultural positioning of models and how they make decisions within complex sociotechnical systems.
    • Computational Reflexivity: Promotes reflexive analysis through scalable data and algorithmic methods, complementing traditional reflexive approaches.
    • Annotator Fingerprinting: Uses data structures to represent annotator behavior patterns, including both human annotators and machine models.
    • Position Mining: Employs unsupervised clustering techniques to identify annotator viewpoint groups and explain predictive biases in the data.
  • Implementation Steps and Key Techniques:

    • Utilized topic modeling methods (e.g., LDA) to decompose annotated data and address annotation sparsity issues.
    • Described annotator and machine learning model behavior patterns using Annotator Fingerprinting.
    • Applied Position Mining techniques (e.g., density-based clustering algorithms like DBSCAN) to identify mainstream or marginal annotation perspectives in the data.

Research Outcomes

  • Specific Findings:

    • Validated the capabilities of Computational Reflexivity and Annotator Fingerprinting through a case study on the Wikipedia Toxic Comments dataset.
    • Identified two major clusters (Cluster 0 and Cluster 1) in the annotated data, representing differing interpretations of "toxicity."
    • Used Annotator Fingerprinting and Position Mining tools to compare and visualize the authors' own positionality with that of other annotators and models.
  • Comparison with Existing Solutions and Advantages:

    • Provides higher-fidelity analytical tools that directly address the limitations of traditional methods in capturing annotator diversity.
    • Encourages data scientists to reflect on how their personal values influence the model development process, while offering more transparent statistical and operational analysis methods.
  • Experimental or Evaluation Results:

    • Leveraged the Empath tool to quantify differences in the distribution of significant word categories in the text, revealing critical disparities among annotators.
    • Demonstrated that data scientists could gain deeper insights into the impact of models within sociotechnical environments by analyzing their own annotation tasks and model positionality.
  • Limitations and Future Directions:

    • While Annotator Fingerprinting addresses annotation sparsity, there is a risk of it being used to obscure inherent biases in research data or methods.
    • Limitations include increased complexity potentially leading to misuse, high technical barriers, and partial interpretations of model behavior.
    • Future work could explore better integration of Computational Reflexivity techniques and language data analysis tools to guide education and methodologies in data science.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/68940/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/abs/10.1145/3491102.3501998
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2022
emoji_events
Award
Honorable Mention
group
Authors
2 authors
sell
Subtopics
Explainable AI (XAI), AI Ethics, Fairness & Accountability, Privacy by Design & User Control
work
Professions
AI/ML Researchers & Engineers, HCI Researchers, Statisticians & Data Scientists
article
Content Status
Full text indexed
hub
Related Papers
6 related papers