"Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AI

Best Paper
Explainable AI (XAI)Algorithmic Fairness & BiasAI/ML Researchers & EngineersHCI Researchers

Title of the Paper

“Everyone wants to do the model work, not the data work”: Data Cascades in High-Stakes AI

Paper Information

  • Domain: Data quality issues and solutions in high-stakes AI applications.
  • Keywords: Data quality, artificial intelligence, machine learning, high-stakes domains, data cascade effects, human-computer interaction, data ethics, data collection, global North-South equity, data documentation.

Research Background and Problem Statement

  • Key Issues Identified:

    • The importance of data quality for AI system performance and downstream impacts is severely underestimated.
    • Data-related tasks are often regarded as “operational” and lack incentives compared to developing new models and algorithms.
    • In high-stakes domains (e.g., healthcare, environmental protection), data quality issues can lead to catastrophic consequences for vulnerable groups and society.
    • Data scarcity and interdisciplinary collaboration challenges further exacerbate the complexity of data quality issues.
  • Significance:

    • Data quality issues in high-stakes domains can result in erroneous predictions, such as cancer misdiagnosis or failed poverty alleviation programs, causing significant societal and individual losses.
    • High-quality data is critical for ensuring fairness, robustness, safety, and scalability in AI systems.
  • Motivation and Related Work:

    • Current AI research often prioritizes model development while neglecting systematic improvements in data quality.
    • The human-computer interaction (HCI) field has begun exploring data packaging, processing, and collaborative work practices, but has not sufficiently addressed data issues in high-stakes domains.

Solutions

  • Methods and Solutions:

    • Defined and introduced the concept of “data cascade effects,” referring to negative chain reactions caused by data issues that may accumulate technical debt over time.
    • Characteristics of data cascade effects: complex, long-term, lack of clear diagnostic indicators, difficult to predict, and negatively impact systems, such as requiring re-iterations or project termination.
    • Proposed structured pathways for improving data quality, emphasizing data documentation, real-time monitoring, and cross-organizational collaboration.
  • Innovations:

    • The study identified specific triggers and manifestations of data issues in high-stakes AI applications.
    • Explored the characteristics of data cascade effects and the interdisciplinary collaboration, data ethics, and standardized documentation methods needed to address these issues.
  • Implementation Steps and Key Techniques:

    • Conducted interviews to understand problems within internal and external data workflows, covering data sources, feedback loops, incentive structures, and upstream-downstream impacts.
    • Proposed standardized metrics for data quality (e.g., fidelity and validity of phenomena) as systematic tools to counter cascade effects.
    • Adjusted data collection practices and collaborated deeply with domain experts to ensure data representativeness and reproducibility.

Research Outcomes

  • Specific Findings:

    • Data cascade effects are widespread in high-stakes domains (92% of respondents experienced at least one data cascade effect).
    • Identified four core data cascade phenomena: instability in the physical world, lack of domain expert knowledge, conflicting reward systems, and missing cross-organizational documentation.
    • Proposed a framework for improving data quality, including educational reforms, tool development, and global data equity initiatives.
  • Advantages and Comparisons:

    • Compared to traditional AI/ML data processing methods, this study systematically highlights the complex impacts of data quality issues in high-stakes domains for the first time.
    • Innovative solutions, such as data-aware tools, standardized documentation, and effective cross-team collaboration practices, can significantly reduce data cascade effects.
  • Experimental and Evaluation Results:

    • Data issues predominantly occur during the early stages of project development (upstream), but their impacts only become apparent later (downstream).
    • By introducing stricter documentation and monitoring measures, some teams successfully mitigated the negative impacts of data issues.
  • Limitations and Future Directions:

    • Limitations:
      • Due to COVID-19 restrictions, the study was unable to conduct in-depth observations of workflows in real-world settings.
      • Gender imbalance in the study reflects the broader gender disparity in the AI industry.
    • Future Directions:
      • Promote the development of interdisciplinary curricula that integrate data quality, ethics, and human-computer interaction into AI education.
      • Develop new tools and interfaces to support data collection, management, and monitoring in high-stakes domains.
      • Advocate for global AI equity frameworks to achieve equal sharing and access to data resources.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/47578/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3411764.3445518
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2021
emoji_events
Award
Best Paper
group
Authors
6 authors
sell
Subtopics
Explainable AI (XAI), Algorithmic Fairness & Bias
work
Professions
AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers