Data-Centric Explanations: Explaining Training Data of Machine Learning Systems to Promote Transparency

Explainable AI (XAI)Privacy by Design & User ControlPrivacy Perception & Decision-MakingCybersecurity EngineersAI/ML Researchers & EngineersHCI Researchers

Title of the Paper

Data-Centric Explanations: Explaining Training Data of Machine Learning Systems to Promote Transparency

Paper Information

  • Research Domain: Human-Computer Interaction (HCI), Machine Learning Transparency and Explainability
  • Keywords: Machine Learning Systems, Dataset Explanation, Transparency, Trust, Fairness, User Experience

Research Background and Problem Statement

  • Problems or Challenges:
    1. The lack of transparency in the training data of machine learning systems may lead to user misunderstandings of system outputs and a loss of trust in the system.
    2. Biases in training data often result in unfairness in system decisions, yet few systems currently explain data-related information to end users.
  • Significance: The performance of machine learning systems heavily depends on the quality and characteristics of training data. Failure to effectively convey this information may negatively impact system transparency, fairness, and user trust, especially in high-stakes decision-making systems.
  • Research Motivation:
    • Explore effective ways to communicate training data-related information to end users.
    • Analyze the impact of such information on users' perceptions of system fairness and trust.
    • Address the limitations of existing explanation systems that focus primarily on outcomes, extending explanations to the data level.

Solution

  • Proposed Approach:
    1. Introduce the concept of "Data-Centric Explanations," which explains the details of machine learning system training data to end users.
    2. Develop a prototype explanation interface based on existing training data documentation methods (e.g., Gebru et al.'s "Datasheets for Datasets") and optimize its usability through user-centered design.
  • Innovations:
    1. Extend traditional explanations focused on system decision-making processes to the dimension of training data.
    2. Conduct in-depth research on user experience, exploring perception differences among users with varying backgrounds regarding data-related explanations.
    3. Propose effective applications of data-centric explanations in different contexts (high-risk vs. low-risk systems).
  • Implementation Steps and Key Techniques:
    1. Identify training data categories suitable for end-user understanding from the literature (including data collection methods, demographic information, intended uses, potential issues, etc.).
    2. Design a user interface presenting information in a Q&A format, refined through preliminary testing and feedback.
    3. Conduct two rounds of user studies to analyze the acceptance and effectiveness of data-centric explanations.

Research Findings

  • Specific Outcomes:
    • Developed a prototype system showcasing five categories of training data information, organized in a Q&A format, including "Data Collection," "Demographics," "Intended Uses," "Potential Issues," and "General Information."
    • Testing demonstrated that data-centric explanations enhance users' perceptions of system transparency, fairness, and trust.
    • User studies highlighted the importance of demographic data (e.g., gender, racial distribution) in helping users identify biases in the data.
  • Advantages Compared to Existing Solutions:
    1. Data-centric explanations are model-agnostic and global, applicable to various machine learning models.
    2. Unlike traditional explanation methods, they reveal the transparency of training data, allowing users to more intuitively identify potential biases.
  • Experimental or Evaluation Results:
    1. Quantitative Survey Results:
      • When explanations revealed balanced training data, users' ratings of system trust, fairness, and comfort significantly increased.
      • No significant differences were observed in perceptions and trust across respondents with different backgrounds (experts, intermediates, beginners).
    2. Qualitative Interview Results: All 27 interviewees agreed that data-centric explanations changed their evaluation of system trust and fairness, especially in high-risk scenarios (e.g., interview decision systems, judicial bail systems).
  • Limitations and Future Directions:
    1. Limitations:
      • The study was conducted in hypothetical scenarios, which may lack ecological validity for real-world systems.
      • Behavioral data analysis of users was not included, with most findings based on subjective feedback.
    2. Potential Improvements:
      • Integrate explanations into real-world systems to obtain ecologically valid data.
      • Combine traditional explanation methods to comprehensively present both data and decision-making processes, enhancing overall transparency.
      • Design more targeted explanation formats for diverse user groups (e.g., decision-makers, media).

Conclusion

This study demonstrates that data-centric explanations significantly enhance user trust and perceptions of fairness in machine learning systems, particularly in high-risk decision-making scenarios. It provides an important direction for advancing AI transparency and fostering societal trust.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/47906/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3411764.3445736
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2021
emoji_events
Award
No award tagged
group
Authors
2 authors
sell
Subtopics
Explainable AI (XAI), Privacy by Design & User Control, Privacy Perception & Decision-Making
work
Professions
Cybersecurity Engineers, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
8 related papers