Quantitative Evaluation of Machine Learning Explanations: A Human-Grounded Benchmark

Explainable AI (XAI)Visualization Perception & CognitionSoftware Engineers & DevelopersAI/ML Researchers & EngineersHCI Researchers

Title of the Paper

Quantitative Evaluation of Machine Learning Explanations: A Human-Grounded Benchmark

Paper Information

  • Research Area: Human-Computer Interaction, Explainable Artificial Intelligence (XAI), Interpretability of Machine Learning Models
  • Keywords: Machine Learning Explanations, Explanation Evaluation, Explanation Benchmark, Data Annotation, Human Attention

Research Background and Problem

  • Identified Problems/Challenges:

    1. Current evaluation methods for explainable machine learning models face a trade-off between objectivity and subjectivity in assessment:
      • Objective evaluations measure the fidelity of explanations to model performance but fail to capture human cognitive agreement with the explanations.
      • Subjective evaluations rely on user feedback, which is costly and prone to subjective bias.
    2. User bias and cognitive limitations reduce the precision of evaluations based on subjective feedback.
    3. There is a lack of a consistent and efficient framework to simultaneously evaluate the correctness and completeness of machine learning model explanations.
  • Significance of the Research:

    • Interpretability is critical for enhancing user trust in machine learning models, especially in high-risk scenarios involving safety, legal, or ethical considerations.
    • Providing a reliable evaluation benchmark can drive the refined design of XAI systems.
  • Research Motivation and Related Work:

    • The authors reviewed common explanation generation methods (e.g., Grad-CAM, LIME) and highlighted their limitations in evaluation standards.
    • Related studies have linked user trust to the "meaningfulness" of explanations, but few have proposed a unified framework integrating human attention and objective evaluation.

Proposed Solution

  • Proposed Method/Solution:

    • The authors propose a benchmark based on multi-layer human attention annotations, which serves as a quantitative evaluation standard for "saliency explanations" of machine learning models.
    • This benchmark collects attention information from multiple annotators to generate multi-layer human attention masks, applicable to both image and text domains.
  • Innovations:

    1. Introduced multi-layer human attention masks to achieve finer-grained feature segmentation of saliency explanations, enabling quantitative analysis of "correctness" (false positives) and "completeness" (false negatives).
    2. Revealed systematic differences between objective and subjective evaluation methods and the impact of user bias through comparative analysis.
    3. Employed a threshold-independent framework using Mean Absolute Error (MAE) as a standard evaluation metric for saliency explanations, avoiding the influence of subjective threshold selection.
  • Implementation Steps/Key Techniques:

    1. Data Annotation Construction:
      • Extracted samples from image datasets (PASCAL VOC and ImageNet) and text datasets (20 Newsgroup and IMDB).
      • Collected attention annotations from 10 unique annotators per sample via the Amazon Mechanical Turk (AMT) platform.
    2. Data Processing:
      • Aggregated annotation results to generate multi-layer masks, removing redundant background pixel interference.
    3. Quantitative Evaluation Experiments:
      • Compared the proposed benchmark with existing single-layer target segmentation mask benchmarks.
      • Validated the correlation between subjective human ratings and the benchmark through user experiments.
    4. Quantification of User Bias:
      • Examined the impact of different error types (FP, FN) and visual presentation styles (e.g., smooth-based Grad-CAM and chunky-style LIME) on subjective ratings.

Research Outcomes

  • Specific Results:

    1. Developed a publicly available human attention benchmark covering both image and text domains.
    2. Demonstrated through experiments that multi-layer human attention masks capture more fine-grained feature saliency compared to single-layer target segmentation masks, aligning more closely with human cognition.
    3. Analyzed discrepancies between human subjective evaluations and the benchmark, revealing manifestations of user rating biases.
    4. User experiments showed:
      • Users tend to rate smoother and more intuitive saliency maps higher (e.g., Grad-CAM received higher ratings compared to LIME).
      • Users are more tolerant of "non-target feature selection errors" (FP) but are more sensitive to "target feature omission errors" (FN).
  • Advantages Compared to Existing Solutions:

    1. Provides a reproducible, threshold-independent quantitative evaluation metric.
    2. Captures more fine-grained information through multi-layer masks, aligning better with subjective user ratings.
    3. Reduces the high cost and evaluation bias associated with multiple rounds of subjective human feedback.
  • Experimental or Evaluation Results:

    • Using quantitative metrics (MAE), the quality of saliency maps generated by image classification models (e.g., Grad-CAM) was compared, supporting the effectiveness of multi-layer human attention masks.
    • Observed that user ratings for certain saliency maps were significantly lower than benchmark evaluations, indicating that the visual appearance of features and error types influence user ratings.
  • Limitations and Future Directions:

    1. Limitations:
      • Data collection and human attention annotation are costly.
      • Experiments primarily involved non-expert participants on crowdsourcing platforms, which may not generalize to domain-specific expert tasks.
    2. Future Directions:
      • Explore universal human attention patterns across different data types and features to further standardize data annotation.
      • Apply the benchmark during the model training phase to optimize model prediction rationality and validate its practical benefits in improving explanation quality.
      • Cross-domain extension: Apply the benchmark to more diverse data types and application scenarios (e.g., medical imaging, legal text).

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/iui/57964/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3397481.3450689
At a Glance

Paper Snapshot

fact_check
dataset
Source
IUI
calendar_month
Year
2021
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Explainable AI (XAI), Visualization Perception & Cognition
work
Professions
Software Engineers & Developers, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
6 related papers