Navigating Uncertainties: How GenAI Developers Document Their Models on Open-Source Platforms
Authors
Paper Title
Navigating Uncertainties: How GenAI Developers Document Their Models on Open-Source Platforms
Publication Info
- Topic area: Documentation practices for Generative AI (GenAI) models in open-source environments.
- Keywords: Generative AI, model documentation, responsible AI, open-source platforms, epistemic uncertainty, methodological uncertainty, ecosystemic uncertainty, model evaluation, AI governance, supply chain accountability.
Background and Problem
- Problem / challenge: Existing AI/ML documentation frameworks are inadequate for addressing the unique challenges posed by GenAI models, which involve open-ended tasks, emergent behaviors, and fragmented governance structures.
- Significance: Effective documentation is critical for ensuring transparency, accountability, and responsible AI development, especially given the risks and societal impacts of GenAI systems.
- Motivation and related work: Previous research has focused on traditional ML documentation in organizational settings and open-source software documentation, but these studies do not account for GenAI’s shifted conditions, such as contested evaluation metrics, open-source dynamics, and fragmented accountability across supply chains.
Solution
- Proposed approach: The study investigates how GenAI developers document their models on open-source platforms and identifies the challenges they face, proposing interventions to address these issues.
- Novelty:
- Introduces "uncertainty" as a conceptual lens to understand GenAI documentation challenges.
- Identifies three forms of uncertainty: normative and epistemic, methodological, and ecosystemic.
- Offers design recommendations for infrastructural, community-based, and collaborative solutions to improve documentation practices.
- Procedure and key techniques:
- Conducted semi-structured interviews with 17 GenAI developers from diverse roles and contexts.
- Analyzed documentation artifacts and interview data using reflexive thematic analysis.
- Identified themes related to uncertainties in documentation practices and proposed actionable recommendations.
Results
- Concrete findings:
- Developers face three forms of uncertainty:
- Normative and epistemic uncertainty: Difficulty determining what to document due to contested norms and lack of measurable model characteristics.
- Methodological uncertainty: Lack of clear procedures for evaluating and documenting model properties, leading to reliance on proxy metrics.
- Ecosystemic uncertainty: Fragmented accountability across the GenAI supply chain, with unclear roles for documenting biases, risks, and appropriate use.
- Developers face three forms of uncertainty:
- Advantage over baselines: Highlights challenges unique to GenAI documentation that are not addressed by traditional ML documentation frameworks, offering tailored solutions for these gaps.
- Experiments / evaluation:
- Interviews with 17 participants from academia, industry, and independent open-source communities.
- Analysis of documentation artifacts, including model cards and usage notes, to triangulate findings.
- Limitations and future work:
- Geographic focus on the US and EU limits global perspectives.
- Small sample size may not capture all documentation practices.
- Future work should include perspectives of documentation consumers, investigate power dynamics within organizations, and explore practices on other platforms.
Summary
This study examines how GenAI developers document their models on open-source platforms, identifying three forms of uncertainty—normative and epistemic, methodological, and ecosystemic—that shape their practices. Developers struggle with determining what to document, how to evaluate and communicate model properties, and who should be responsible for documentation across the supply chain. The findings highlight the need for infrastructural support, community-driven norm-building, and collaborative frameworks to address these challenges. By addressing these uncertainties, the study aims to improve transparency and accountability in the rapidly evolving GenAI ecosystem.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
The Disagreement Deconvolution: Bringing Machine Learning Performance Metrics In Line With Reality
CHI '21· Explainable AI (XAI) +1
- 71%
HILL: A Hallucination Identifier for Large Language Models
CHI '24· Explainable AI (XAI) +2
- 63%
Evaluating the Interpretability of Generative Models by Interactive Reconstruction
CHI '21· Explainable AI (XAI) +2
- 63%
Unraveling the Dilemma of AI Errors: Exploring the Effectiveness of Human and Machine Explanations for Large Language Models
CHI '24· Human-LLM Collaboration +2
- 63%
Access Denied: Meaningful Data Access for Quantitative Algorithm Audits
CHI '25· Explainable AI (XAI) +2
- 63%
Towards AI Accountability Infrastructure: Gaps and Opportunities in AI Audit Tooling
CHI '25· Explainable AI (XAI) +2
- 63%
Do People Appropriately Rely on AI-Advice? An Analytical Review of HCI Research on Human-AI Decision-Making
CHI '26· AI-Assisted Decision-Making & Automation +2
- 63%
Red Teaming LLMs as Socio-Technical Practice: From Exploration and Data Creation to Evaluation
CHI '26· Explainable AI (XAI) +2
Based on Jaccard similarity of research subtopics & professions (≥60%)