Data-Centric Explanations: Explaining Training Data of Machine Learning Systems to Promote Transparency
Title of the Paper
Data-Centric Explanations: Explaining Training Data of Machine Learning Systems to Promote Transparency
Paper Information
- Research Domain: Human-Computer Interaction (HCI), Machine Learning Transparency and Explainability
- Keywords: Machine Learning Systems, Dataset Explanation, Transparency, Trust, Fairness, User Experience
Research Background and Problem Statement
- Problems or Challenges:
- The lack of transparency in the training data of machine learning systems may lead to user misunderstandings of system outputs and a loss of trust in the system.
- Biases in training data often result in unfairness in system decisions, yet few systems currently explain data-related information to end users.
- Significance: The performance of machine learning systems heavily depends on the quality and characteristics of training data. Failure to effectively convey this information may negatively impact system transparency, fairness, and user trust, especially in high-stakes decision-making systems.
- Research Motivation:
- Explore effective ways to communicate training data-related information to end users.
- Analyze the impact of such information on users' perceptions of system fairness and trust.
- Address the limitations of existing explanation systems that focus primarily on outcomes, extending explanations to the data level.
Solution
- Proposed Approach:
- Introduce the concept of "Data-Centric Explanations," which explains the details of machine learning system training data to end users.
- Develop a prototype explanation interface based on existing training data documentation methods (e.g., Gebru et al.'s "Datasheets for Datasets") and optimize its usability through user-centered design.
- Innovations:
- Extend traditional explanations focused on system decision-making processes to the dimension of training data.
- Conduct in-depth research on user experience, exploring perception differences among users with varying backgrounds regarding data-related explanations.
- Propose effective applications of data-centric explanations in different contexts (high-risk vs. low-risk systems).
- Implementation Steps and Key Techniques:
- Identify training data categories suitable for end-user understanding from the literature (including data collection methods, demographic information, intended uses, potential issues, etc.).
- Design a user interface presenting information in a Q&A format, refined through preliminary testing and feedback.
- Conduct two rounds of user studies to analyze the acceptance and effectiveness of data-centric explanations.
Research Findings
- Specific Outcomes:
- Developed a prototype system showcasing five categories of training data information, organized in a Q&A format, including "Data Collection," "Demographics," "Intended Uses," "Potential Issues," and "General Information."
- Testing demonstrated that data-centric explanations enhance users' perceptions of system transparency, fairness, and trust.
- User studies highlighted the importance of demographic data (e.g., gender, racial distribution) in helping users identify biases in the data.
- Advantages Compared to Existing Solutions:
- Data-centric explanations are model-agnostic and global, applicable to various machine learning models.
- Unlike traditional explanation methods, they reveal the transparency of training data, allowing users to more intuitively identify potential biases.
- Experimental or Evaluation Results:
- Quantitative Survey Results:
- When explanations revealed balanced training data, users' ratings of system trust, fairness, and comfort significantly increased.
- No significant differences were observed in perceptions and trust across respondents with different backgrounds (experts, intermediates, beginners).
- Qualitative Interview Results: All 27 interviewees agreed that data-centric explanations changed their evaluation of system trust and fairness, especially in high-risk scenarios (e.g., interview decision systems, judicial bail systems).
- Quantitative Survey Results:
- Limitations and Future Directions:
- Limitations:
- The study was conducted in hypothetical scenarios, which may lack ecological validity for real-world systems.
- Behavioral data analysis of users was not included, with most findings based on subjective feedback.
- Potential Improvements:
- Integrate explanations into real-world systems to obtain ecologically valid data.
- Combine traditional explanation methods to comprehensively present both data and decision-making processes, enhancing overall transparency.
- Design more targeted explanation formats for diverse user groups (e.g., decision-makers, media).
- Limitations:
Conclusion
This study demonstrates that data-centric explanations significantly enhance user trust and perceptions of fairness in machine learning systems, particularly in high-risk decision-making scenarios. It provides an important direction for advancing AI transparency and fostering societal trust.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can training data of machine learning systems be explained to users to improve transparency?Category: Machine Learning Fairness and Data Development PracticesSimilar questionsarrow_forward
- How do data-level explanations affect users' perceptions of system fairness and trustworthiness?Category: Machine Learning Fairness and Data Development PracticesSimilar questionsarrow_forward
- Which categories of training data are suitable to communicate to users to help identify potential bias?Category: Machine Learning Fairness and Data Development PracticesSimilar questionsarrow_forward
Practical Problems
1- Users cannot understand machine learning systems' training data sources and potential bias, making it difficult to trust system decisions.Category: Machine Learning Fairness and Data Development PracticesSimilar questionsarrow_forward
- 71%
Rescriber: Smaller-LLM-Powered User-Led Data Minimization for LLM-Based Chatbots
CHI '25· Explainable AI (XAI) +2
- 71%
"Having Confidence in My Confidence Intervals": How Data Users Engage with Privacy-Protected Wikipedia Data
CHI '26· Explainable AI (XAI) +2
- 71%
How Much Trust is Enough? Towards Calibrating Trust in Technology
CHI '26· Explainable AI (XAI) +2
- 71%
When the Codec Hallucinates: User Perceptions of Miscompressed Images
CHI '26· Explainable AI (XAI) +2
- 67%
Don’t Just Tell Me, Ask Me: AI Systems that Intelligently Frame Explanations as Questions Improve Human Logical Discernment Accuracy over Causal AI explanations
CHI '23· Explainable AI (XAI) +1
- 63%
What is Sensitive About (Sensitive) Data? Characterizing Sensitivity and Intimacy of Google Assistant Speech Records
CHI '23· Explainable AI (XAI) +3
- 63%
“Technically speaking I’m at the top of the hierarchy”: How System Administrators Think About Power
CHI '26· Explainable AI (XAI) +3
- 63%
Tool-Assisted CVSS Vulnerability Scoring: A Controlled Quantitative Study of Human Assessment
CHI '26· Explainable AI (XAI) +3
Based on Jaccard similarity of research subtopics & professions (≥60%)