"Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AI
Best PaperAuthors
Title of the Paper
“Everyone wants to do the model work, not the data work”: Data Cascades in High-Stakes AI
Paper Information
- Domain: Data quality issues and solutions in high-stakes AI applications.
- Keywords: Data quality, artificial intelligence, machine learning, high-stakes domains, data cascade effects, human-computer interaction, data ethics, data collection, global North-South equity, data documentation.
Research Background and Problem Statement
-
Key Issues Identified:
- The importance of data quality for AI system performance and downstream impacts is severely underestimated.
- Data-related tasks are often regarded as “operational” and lack incentives compared to developing new models and algorithms.
- In high-stakes domains (e.g., healthcare, environmental protection), data quality issues can lead to catastrophic consequences for vulnerable groups and society.
- Data scarcity and interdisciplinary collaboration challenges further exacerbate the complexity of data quality issues.
-
Significance:
- Data quality issues in high-stakes domains can result in erroneous predictions, such as cancer misdiagnosis or failed poverty alleviation programs, causing significant societal and individual losses.
- High-quality data is critical for ensuring fairness, robustness, safety, and scalability in AI systems.
-
Motivation and Related Work:
- Current AI research often prioritizes model development while neglecting systematic improvements in data quality.
- The human-computer interaction (HCI) field has begun exploring data packaging, processing, and collaborative work practices, but has not sufficiently addressed data issues in high-stakes domains.
Solutions
-
Methods and Solutions:
- Defined and introduced the concept of “data cascade effects,” referring to negative chain reactions caused by data issues that may accumulate technical debt over time.
- Characteristics of data cascade effects: complex, long-term, lack of clear diagnostic indicators, difficult to predict, and negatively impact systems, such as requiring re-iterations or project termination.
- Proposed structured pathways for improving data quality, emphasizing data documentation, real-time monitoring, and cross-organizational collaboration.
-
Innovations:
- The study identified specific triggers and manifestations of data issues in high-stakes AI applications.
- Explored the characteristics of data cascade effects and the interdisciplinary collaboration, data ethics, and standardized documentation methods needed to address these issues.
-
Implementation Steps and Key Techniques:
- Conducted interviews to understand problems within internal and external data workflows, covering data sources, feedback loops, incentive structures, and upstream-downstream impacts.
- Proposed standardized metrics for data quality (e.g., fidelity and validity of phenomena) as systematic tools to counter cascade effects.
- Adjusted data collection practices and collaborated deeply with domain experts to ensure data representativeness and reproducibility.
Research Outcomes
-
Specific Findings:
- Data cascade effects are widespread in high-stakes domains (92% of respondents experienced at least one data cascade effect).
- Identified four core data cascade phenomena: instability in the physical world, lack of domain expert knowledge, conflicting reward systems, and missing cross-organizational documentation.
- Proposed a framework for improving data quality, including educational reforms, tool development, and global data equity initiatives.
-
Advantages and Comparisons:
- Compared to traditional AI/ML data processing methods, this study systematically highlights the complex impacts of data quality issues in high-stakes domains for the first time.
- Innovative solutions, such as data-aware tools, standardized documentation, and effective cross-team collaboration practices, can significantly reduce data cascade effects.
-
Experimental and Evaluation Results:
- Data issues predominantly occur during the early stages of project development (upstream), but their impacts only become apparent later (downstream).
- By introducing stricter documentation and monitoring measures, some teams successfully mitigated the negative impacts of data issues.
-
Limitations and Future Directions:
- Limitations:
- Due to COVID-19 restrictions, the study was unable to conduct in-depth observations of workflows in real-world settings.
- Gender imbalance in the study reflects the broader gender disparity in the AI industry.
- Future Directions:
- Promote the development of interdisciplinary curricula that integrate data quality, ethics, and human-computer interaction into AI education.
- Develop new tools and interfaces to support data collection, management, and monitoring in high-stakes domains.
- Advocate for global AI equity frameworks to achieve equal sharing and access to data resources.
- Limitations:
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- What are the triggers and characteristics of data cascade effects (chain reactions caused by data quality problems)?Category: Human-in-the-Loop Labeling and Example SelectionSimilar questionsarrow_forward
- How can cross-organizational collaboration and standardized documentation reduce data cascade effects in high-risk domains (e.g., healthcare, environmental protection)?Category: Human-in-the-Loop Labeling and Example SelectionSimilar questionsarrow_forward
- How can data quality metrics be defined and measured in AI development to better address data problems?Category: Human-in-the-Loop Labeling and Example SelectionSimilar questionsarrow_forward
Practical Problems
1- AI in high-risk domains struggles to avoid erroneous predictions caused by data quality problems.Category: Human-in-the-Loop Labeling and Example SelectionSimilar questionsarrow_forward
- 100%
Fairness Evaluation in Text Classification: Machine Learning Practitioner Perspectives of Individual and Group Fairness
CHI '23· Explainable AI (XAI) +1
- 100%
Perceptions of the Fairness Impacts of Multiplicity in Machine Learning
CHI '25· Explainable AI (XAI) +1
- 80%
Beyond Expertise and Roles: A Framework to Characterize the Stakeholders of Interpretable Machine Learning and their Needs
CHI '21· Explainable AI (XAI) +2
- 80%
User-Driven Value Alignment: Understanding Users' Perceptions and Strategies for Addressing Biased and Discriminatory Statements in AI Companions
CHI '25· Explainable AI (XAI) +2
- 80%
More than Marketing? On the Information Value of AI Benchmarks for Practitioners
IUI '25· Explainable AI (XAI) +1
- 75%
Designing Interactive Explainable AI Tools for Algorithmic Literacy and Transparency
DIS '24· Explainable AI (XAI) +1
- 75%
Conversational Explanations: Discussing Explainable AI with Non-AI Experts
IUI '25· Explainable AI (XAI)
- 67%
Do Expressions Change Decisions? Exploring the Impact of AI's Explanation Tone on Decision-Making
CHI '25· Explainable AI (XAI) +2
- 67%
Explaining Models: An Empirical Study of How Explanations Impact Fairness Judgment
IUI '19· Explainable AI (XAI) +2
- 67%
The Impact of Explanations on Fairness in Human-AI Decision-Making: Protected vs Proxy Features
IUI '24· Explainable AI (XAI) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)