Model Positionality and Computational Reflexivity: Promoting Reflexivity in Data Science
Honorable MentionTitle of the Paper
Model Positionality and Computational Reflexivity: Promoting Reflexivity in Data Science
Bibliographic Information
- Subject Areas: Human-Centered Data Science, Machine Learning Ethics, Annotated Data Processing
- Keywords: Computational Reflexivity, Model Positionality, Data Science, Human-Centered Machine Learning, Annotator Fingerprinting, Position Mining, Critical Data Studies, Online Communities
- Authors and Emails: Scott Allen Cambo (scottallencambo@gmail.com), Darren Gergle (dgergle@northwestern.edu)
- Publication Venue and Date: CHI Conference on Human Factors in Computing Systems (CHI ’22), April 29-May 5, 2022
- DOI: https://doi.org/10.1145/3491102.3501998
Research Background and Problem
-
Problems or Challenges Identified by the Authors:
- Although data science and machine learning methods aim for objective analysis, subjective decisions during their development and deployment often go unacknowledged, leading to potential biases.
- Models used on social platforms, such as algorithms for detecting online harassment, face practical limitations due to subjective human interpretations of concepts like "toxicity."
- Large-scale machine learning workflows lack sufficient transparency regarding the sociocultural backgrounds of data scientists, annotators, and models.
-
Why This Problem is Important:
- Data science is widely applied in critical areas affecting human lives (e.g., criminal justice, vaccine distribution), where potential biases and erroneous decisions can have profound implications for social justice.
- A lack of deep reflection on the model development process can result in unfair treatment of minority groups and even systemic oppression.
-
Research Motivation and Related Work:
- Building on the trend of human-centered data science, the authors aim to promote transparency and reflexive analysis in decision-making processes.
- Drawing inspiration from qualitative research methods in social sciences, such as positionality and reflexivity, the authors extend these concepts to the context of data science.
Solution
-
Proposed Solution by the Authors:
- Introduced two core concepts: "Model Positionality" and "Computational Reflexivity."
- Developed new techniques: Annotator Fingerprinting and Position Mining, to help understand the sociocultural backgrounds of models, annotators, and data scientists.
-
Innovative Aspects of the Solution:
- Model Positionality: Reveals the social and cultural positioning of models and how they make decisions within complex sociotechnical systems.
- Computational Reflexivity: Promotes reflexive analysis through scalable data and algorithmic methods, complementing traditional reflexive approaches.
- Annotator Fingerprinting: Uses data structures to represent annotator behavior patterns, including both human annotators and machine models.
- Position Mining: Employs unsupervised clustering techniques to identify annotator viewpoint groups and explain predictive biases in the data.
-
Implementation Steps and Key Techniques:
- Utilized topic modeling methods (e.g., LDA) to decompose annotated data and address annotation sparsity issues.
- Described annotator and machine learning model behavior patterns using Annotator Fingerprinting.
- Applied Position Mining techniques (e.g., density-based clustering algorithms like DBSCAN) to identify mainstream or marginal annotation perspectives in the data.
Research Outcomes
-
Specific Findings:
- Validated the capabilities of Computational Reflexivity and Annotator Fingerprinting through a case study on the Wikipedia Toxic Comments dataset.
- Identified two major clusters (Cluster 0 and Cluster 1) in the annotated data, representing differing interpretations of "toxicity."
- Used Annotator Fingerprinting and Position Mining tools to compare and visualize the authors' own positionality with that of other annotators and models.
-
Comparison with Existing Solutions and Advantages:
- Provides higher-fidelity analytical tools that directly address the limitations of traditional methods in capturing annotator diversity.
- Encourages data scientists to reflect on how their personal values influence the model development process, while offering more transparent statistical and operational analysis methods.
-
Experimental or Evaluation Results:
- Leveraged the Empath tool to quantify differences in the distribution of significant word categories in the text, revealing critical disparities among annotators.
- Demonstrated that data scientists could gain deeper insights into the impact of models within sociotechnical environments by analyzing their own annotation tasks and model positionality.
-
Limitations and Future Directions:
- While Annotator Fingerprinting addresses annotation sparsity, there is a risk of it being used to obscure inherent biases in research data or methods.
- Limitations include increased complexity potentially leading to misuse, high technical barriers, and partial interpretations of model behavior.
- Future work could explore better integration of Computational Reflexivity techniques and language data analysis tools to guide education and methodologies in data science.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can model situating reveal the social and cultural contexts of data models?Category: Machine Learning Fairness and Data Development PracticesSimilar questionsarrow_forward
- How does computational reflexivity support reflexive analysis in data science?Category: Machine Learning Fairness and Data Development PracticesSimilar questionsarrow_forward
- How can annotator fingerprints and stance mining be used to analyze bias in models and annotators?Category: Machine Learning Fairness and Data Development PracticesSimilar questionsarrow_forward
Practical Problems
1- Subjective decisions in data science model development easily introduce bias and affect fair decision-making.Category: Machine Learning Fairness and Data Development PracticesSimilar questionsarrow_forward
- 71%
Decomposing Autonomy: Explaining AI Technology Acceptance Through a Liberty-Based Framework
CHI '26· Explainable AI (XAI) +2
- 71%
Certified AI System = Trustworthy? Exploring Expert and Lay User Perceptions and Needs Regarding AI Certification
CHI '26· Explainable AI (XAI) +2
- 67%
Don’t Just Tell Me, Ask Me: AI Systems that Intelligently Frame Explanations as Questions Improve Human Logical Discernment Accuracy over Causal AI explanations
CHI '23· Explainable AI (XAI) +1
- 63%
PASTA: A Scalable Framework for Multi-Policy AI Compliance Evaluation
CHI '26· Explainable AI (XAI) +3
- 63%
Characterizing User-Reported Risks across LLM Chatbots
CHI '26· Human-LLM Collaboration +3
- 63%
AI and My Values: User Perceptions of LLMs’ Ability to Extract, Embody, and Explain Human Values from Casual Conversations
CHI '26· Human-LLM Collaboration +3
Based on Jaccard similarity of research subtopics & professions (≥60%)