Quantitative Evaluation of Machine Learning Explanations: A Human-Grounded Benchmark
Authors
Explainable AI (XAI)Visualization Perception & CognitionSoftware Engineers & DevelopersAI/ML Researchers & EngineersHCI Researchers
Title of the Paper
Quantitative Evaluation of Machine Learning Explanations: A Human-Grounded Benchmark
Paper Information
- Research Area: Human-Computer Interaction, Explainable Artificial Intelligence (XAI), Interpretability of Machine Learning Models
- Keywords: Machine Learning Explanations, Explanation Evaluation, Explanation Benchmark, Data Annotation, Human Attention
Research Background and Problem
-
Identified Problems/Challenges:
- Current evaluation methods for explainable machine learning models face a trade-off between objectivity and subjectivity in assessment:
- Objective evaluations measure the fidelity of explanations to model performance but fail to capture human cognitive agreement with the explanations.
- Subjective evaluations rely on user feedback, which is costly and prone to subjective bias.
- User bias and cognitive limitations reduce the precision of evaluations based on subjective feedback.
- There is a lack of a consistent and efficient framework to simultaneously evaluate the correctness and completeness of machine learning model explanations.
- Current evaluation methods for explainable machine learning models face a trade-off between objectivity and subjectivity in assessment:
-
Significance of the Research:
- Interpretability is critical for enhancing user trust in machine learning models, especially in high-risk scenarios involving safety, legal, or ethical considerations.
- Providing a reliable evaluation benchmark can drive the refined design of XAI systems.
-
Research Motivation and Related Work:
- The authors reviewed common explanation generation methods (e.g., Grad-CAM, LIME) and highlighted their limitations in evaluation standards.
- Related studies have linked user trust to the "meaningfulness" of explanations, but few have proposed a unified framework integrating human attention and objective evaluation.
Proposed Solution
-
Proposed Method/Solution:
- The authors propose a benchmark based on multi-layer human attention annotations, which serves as a quantitative evaluation standard for "saliency explanations" of machine learning models.
- This benchmark collects attention information from multiple annotators to generate multi-layer human attention masks, applicable to both image and text domains.
-
Innovations:
- Introduced multi-layer human attention masks to achieve finer-grained feature segmentation of saliency explanations, enabling quantitative analysis of "correctness" (false positives) and "completeness" (false negatives).
- Revealed systematic differences between objective and subjective evaluation methods and the impact of user bias through comparative analysis.
- Employed a threshold-independent framework using Mean Absolute Error (MAE) as a standard evaluation metric for saliency explanations, avoiding the influence of subjective threshold selection.
-
Implementation Steps/Key Techniques:
- Data Annotation Construction:
- Extracted samples from image datasets (PASCAL VOC and ImageNet) and text datasets (20 Newsgroup and IMDB).
- Collected attention annotations from 10 unique annotators per sample via the Amazon Mechanical Turk (AMT) platform.
- Data Processing:
- Aggregated annotation results to generate multi-layer masks, removing redundant background pixel interference.
- Quantitative Evaluation Experiments:
- Compared the proposed benchmark with existing single-layer target segmentation mask benchmarks.
- Validated the correlation between subjective human ratings and the benchmark through user experiments.
- Quantification of User Bias:
- Examined the impact of different error types (FP, FN) and visual presentation styles (e.g., smooth-based Grad-CAM and chunky-style LIME) on subjective ratings.
- Data Annotation Construction:
Research Outcomes
-
Specific Results:
- Developed a publicly available human attention benchmark covering both image and text domains.
- Demonstrated through experiments that multi-layer human attention masks capture more fine-grained feature saliency compared to single-layer target segmentation masks, aligning more closely with human cognition.
- Analyzed discrepancies between human subjective evaluations and the benchmark, revealing manifestations of user rating biases.
- User experiments showed:
- Users tend to rate smoother and more intuitive saliency maps higher (e.g., Grad-CAM received higher ratings compared to LIME).
- Users are more tolerant of "non-target feature selection errors" (FP) but are more sensitive to "target feature omission errors" (FN).
-
Advantages Compared to Existing Solutions:
- Provides a reproducible, threshold-independent quantitative evaluation metric.
- Captures more fine-grained information through multi-layer masks, aligning better with subjective user ratings.
- Reduces the high cost and evaluation bias associated with multiple rounds of subjective human feedback.
-
Experimental or Evaluation Results:
- Using quantitative metrics (MAE), the quality of saliency maps generated by image classification models (e.g., Grad-CAM) was compared, supporting the effectiveness of multi-layer human attention masks.
- Observed that user ratings for certain saliency maps were significantly lower than benchmark evaluations, indicating that the visual appearance of features and error types influence user ratings.
-
Limitations and Future Directions:
- Limitations:
- Data collection and human attention annotation are costly.
- Experiments primarily involved non-expert participants on crowdsourcing platforms, which may not generalize to domain-specific expert tasks.
- Future Directions:
- Explore universal human attention patterns across different data types and features to further standardize data annotation.
- Apply the benchmark during the model training phase to optimize model prediction rationality and validate its practical benefits in improving explanation quality.
- Cross-domain extension: Apply the benchmark to more diverse data types and application scenarios (e.g., medical imaging, legal text).
- Limitations:
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How can the correctness and completeness of saliency explanations for machine learning models be quantitatively evaluated?Category: XAI Explanation and Appropriate Reliance CalibrationSimilar questionsarrow_forward
- What biases exist in existing explanation evaluation methods based on subjective user ratings?Category: XAI Explanation and Appropriate Reliance CalibrationSimilar questionsarrow_forward
- Can multi-level human attention masks better capture feature saliency consistent with human cognition?Category: XAI Explanation and Appropriate Reliance CalibrationSimilar questionsarrow_forward
lightbulb
Practical Problems
1- Users struggle to trust explanations provided by machine learning models, especially in critical scenarios.Category: XAI Explanation and Appropriate Reliance CalibrationSimilar questionsarrow_forward
- 80%
Graphical Perception of Saliency-based Model Explanations
CHI '23· Explainable AI (XAI) +1
- 67%
The Disagreement Deconvolution: Bringing Machine Learning Performance Metrics In Line With Reality
CHI '21· Explainable AI (XAI) +1
- 67%
"Are You Really Sure?'' Understanding the Effects of Human Self-Confidence Calibration in AI-Assisted Decision Making
CHI '24· Explainable AI (XAI) +1
- 67%
Wikipedia ORES Explorer: Visualizing Trade-offs For Designing Applications With Machine Learning API
DIS '21· Explainable AI (XAI) +1
- 67%
More than Marketing? On the Information Value of AI Benchmarks for Practitioners
IUI '25· Explainable AI (XAI) +1
- 60%
Conversational Explanations: Discussing Explainable AI with Non-AI Experts
IUI '25· Explainable AI (XAI)
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3397481.3450689
At a Glance
fact_checkPaper Snapshot
dataset
Source
IUI
calendar_month
Year
2021
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Explainable AI (XAI), Visualization Perception & Cognition
work
Professions
Software Engineers & Developers, AI/ML Researchers & Engineers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
6 related papers