Treading the Transparency Tightrope: A Taxonomy of Risks and Benefits of Foundation Model Data Transparency for Transparency Advocates

Explainable AI (XAI)Algorithmic Transparency & AuditabilityPrivacy by Design & User ControlPrivacy Policy MakersAI/ML Researchers & EngineersHCI Researchers

Paper Title

Treading the Transparency Tightrope: A Taxonomy of Risks and Benefits of Foundation Model Data Transparency for Transparency Advocates

Publication Info

  • Topic area: Risks and benefits of data transparency in foundation models (FMs).
  • Keywords: AI data transparency, foundation models, dataset documentation, dataset availability, transparency risks, transparency benefits, stakeholder analysis, responsible AI, situational transparency, transparency taxonomy.

Background and Problem

  • Problem / challenge: The opacity of datasets used in foundation models (FMs) impedes independent analysis of harms, accountability, and innovation. Current transparency demands often lack clarity on specific benefits or risks, leading to critiques of their practical relevance.
  • Significance: Transparency is critical for addressing ethical, legal, and societal concerns in AI, such as privacy breaches, intellectual property violations, and fairness issues. However, uncritical transparency can also introduce risks, such as safety concerns and competitive disadvantages.
  • Motivation and related work: Prior work has highlighted the need for transparency in AI systems but often treats transparency as inherently good without considering its contingent nature. This paper builds on critical transparency scholarship and stakeholder analysis to address this gap.

Solution

  • Proposed approach: A taxonomy of risks and benefits of AI data transparency, emphasizing situational transparency that accounts for stakeholder positions and transparency modalities.
  • Novelty:
    1. Identification of four risks (contamination, competitiveness, safety, scrutiny) and four benefits (accountability, innovation, integrity, suitability) of data transparency.
    2. Introduction of two key factors influencing transparency perceptions: stakeholder positions (advocacy vs. opposition) and transparency modalities (data documentation vs. data availability).
    3. Conceptualization of situational data transparency to balance risks and benefits.
    4. Practical considerations for transparency advocates to optimize benefits and mitigate risks.
  • Procedure and key techniques:
    1. Systematic sampling of 153 sources, including academic papers, NGO reports, legal frameworks, and FM developer documentation.
    2. Qualitative coding to identify risks and benefits of transparency, followed by thematic clustering into higher-level categories.
    3. Analysis of the interplay between stakeholder positions, transparency modalities, and perceived risks/benefits.

Results

  • Concrete findings:
    • Risks: Contamination (5.2% of sources), competitiveness (23.5%), safety (39.2%), scrutiny (17.6%).
    • Benefits: Accountability (35.9%), innovation (34%), integrity (47.1%), suitability (9.2%).
    • Risks and benefits are influenced by stakeholder positions and transparency modalities.
  • Advantage over baselines: Provides a nuanced framework for understanding transparency as contingent, rather than inherently good or bad, enabling more strategic advocacy.
  • Experiments / evaluation: Analysis of diverse stakeholder documentation, including researchers, NGOs, law and policy leaders, and FM developers. Coding and thematic analysis were used to derive the taxonomy.
  • Limitations and future work:
    • Limited representation of underrepresented stakeholders, such as data subjects and community developers.
    • Focus on English-language sources and Anglophone contexts.
    • Future research should empirically validate the taxonomy, explore additional risks/benefits, and examine evolving regulatory impacts.

Summary

This paper develops a taxonomy of four risks (contamination, competitiveness, safety, scrutiny) and four benefits (accountability, innovation, integrity, suitability) associated with AI data transparency. It introduces two key factors—stakeholder positions (advocacy vs. opposition) and transparency modalities (data documentation vs. data availability)—to contextualize transparency as contingent. The taxonomy provides a framework for situational data transparency, enabling advocates to balance benefits and risks based on specific goals and contexts. This work offers practical guidance for transparency advocates and lays the groundwork for future research on responsible AI data practices.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223313/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790364
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Explainable AI (XAI), Algorithmic Transparency & Auditability, Privacy by Design & User Control
work
Professions
Privacy Policy Makers, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers