Contextualizing Datasets: Deepening Awareness of Data Biases through Critical Reflection of Civic Data
Paper Title
Contextualizing Datasets: Deepening Awareness of Data Biases through Critical Reflection of Civic Data
Publication Info
- Topic area: Critical data literacy and bias awareness in AI datasets.
- Keywords: Critical data literacy, AI ethics, data bias, civic data, stakeholder perspectives, data contextualization, educational tools, participatory AI, dataset exploration, responsible AI.
Background and Problem
- Problem / challenge: Datasets used in AI are often treated as neutral, despite being shaped by human subjectivity and biases. Non-technical stakeholders lack the tools to critically explore and contextualize datasets, limiting their ability to identify biases and shortcomings.
- Significance: Addressing dataset biases is crucial for ethical AI design, especially in high-stakes domains like public resource allocation, where biased data can perpetuate inequities.
- Motivation and related work: Prior research has highlighted the importance of critical data literacy and engaging diverse stakeholders in AI design. However, existing methods often fail to connect data exploration with actionable insights or to support non-technical users in reflecting on data contexts and biases.
Solution
- Proposed approach: A web-tool called "Contextualizing Datasets" that guides users through data exploration, stakeholder interactions, and critical reflections to deepen awareness of dataset biases.
- Novelty:
- Introduces a structured process for data contextualization, including question cards and stakeholder dialogues.
- Focuses on non-technical users, enabling critical reflection without requiring advanced technical skills.
- Uses a local civic dataset and case study to connect data literacy with lived experiences.
- Provides a framework for adapting and re-purposing the tool for other datasets and contexts.
- Procedure and key techniques:
- Users explore datasets using question cards adapted from data documentation resources.
- Stakeholder dialogues simulate interactions with dataset owners, data workers, and community representatives.
- Critical reflections prompt users to synthesize insights and propose actionable recommendations.
- A case study on Austin’s 311 data for flood resource allocation serves as the testbed for the tool.
Results
- Concrete findings:
- Students’ average support for using 311 data decreased from 5.48 to 5.12 (on a 7-point scale), while concern increased from 4.68 to 5.04.
- 46% of students recommended questions about stakeholder roles and representation as essential for dataset exploration.
- Students identified issues such as data reliability, demographic biases, and the mismatch between the dataset’s original purpose and its current use.
- Advantage over baselines:
- Unlike traditional tools like Excel or Tableau, Contextualizing Datasets explicitly prompts users to consider data contexts and biases.
- Students reported that the tool provided a structured and reflective approach to exploring datasets, which is often missing in standard data analysis workflows.
- Experiments / evaluation:
- Conducted with 25 graduate students in a Human-AI Interaction course.
- Mixed-methods analysis included reflections, question card responses, post-study surveys, and interviews.
- Participants explored Austin’s 311 dataset to prioritize neighborhoods for flood mitigation, guided by the tool.
- Limitations and future work:
- Limited to a single case study with graduate students, who may have higher baseline awareness of data and AI issues.
- Time constraints (one class session) may have limited the depth of exploration.
- Future work includes integrating the tool with AI design activities, expanding stakeholder roles, and adapting it for generative AI contexts.
Summary
This study introduces "Contextualizing Datasets," a web-tool designed to enhance critical data literacy by guiding users through dataset exploration, stakeholder dialogues, and reflective exercises. Applied to a case study on Austin’s 311 data for flood resource allocation, the tool helped graduate students identify biases and deepen their understanding of data contexts. Students appreciated the structured approach and stakeholder perspectives but faced challenges in proposing actionable solutions. The tool shows promise for fostering ethical AI practices and can be adapted for broader educational and professional contexts, including generative AI applications.
Research Questions / Practical Problems
Question signals indexed for this paper.
Based on Jaccard similarity of research subtopics & professions (≥60%)