Towards Relatable Explainable AI with the Perceptual Process
Best PaperEye Tracking & Gaze InteractionExplainable AI (XAI)
Document Title
Towards Relatable Explainable AI with the Perceptual Process
Document Information
- Subject Area: Explainable Artificial Intelligence (XAI), particularly contrastive explanations in perceptual tasks
- Keywords: Explainable AI (XAI), contrastive explanations, speech emotion recognition, deep learning models, user needs, human cognition
Research Background and Issues
-
Problems and Challenges:
- Machine learning models are often highly complex, making them difficult for end users to understand, which limits their real-world applications.
- Current contrastive explanation methods remain overly basic, lack semantic meaning, and are difficult to align with human cognition.
- In the audio data domain, existing explanation methods (e.g., saliency maps based on spectrograms) are overly technical and challenging for general users to comprehend.
-
Significance:
- Explainability is a key factor in achieving responsible and trustworthy artificial intelligence.
- Applications in the audio domain (e.g., speech emotion recognition) require explanation methods that align with human perceptual processes to enhance usability and trust.
-
Research Motivation:
- Inspired by the theory of human perceptual processes, this study proposes a more relatable explanation framework that aligns explanations with users' conceptual, hypothetical, and associative cognitive needs.
- It explores various contrastive explanations, including saliency, counterfactuals, and supplementary cues.
Solution
-
Proposed Framework and Model:
- XAI Perceptual Processing Framework: Based on the theory of human perceptual processes, this framework is divided into three stages: Select, Organize, and Interpret.
- RexNet Model (Relatable Explanation Network): A modular multi-task deep learning framework supporting three types of contrastive explanations:
- Contrastive Saliency: Generates contrastive saliency maps to help users understand the input features the model focuses on.
- Counterfactual Synthetic: Uses a generative adversarial network (StarGAN-VC) to generate style-transformed speech samples, showing differences from the target classification.
- Contrastive Cues: Provides emotion-related speech cues (e.g., high-frequency energy, volume) for comparison.
-
Innovations:
- Incorporates human cognitive psychology by introducing perceptual theories into AI explanation design.
- Proposes a model architecture that organically integrates saliency, counterfactual examples, and cue information.
- First application of relatable explanation methods in audio prediction tasks.
-
Implementation Steps:
- Train a speech emotion classification model using the RAVDESS dataset and extend it into RexNet, achieving emotion prediction through a spectrogram-based CNN.
- Integrate explanation modules and generate explanations using techniques such as Grad-CAM, generative adversarial networks (GAN), and layer-wise relevance propagation (LRP).
- Evaluate the usability and effectiveness of the explanations through modeling studies and user experiments.
Research Outcomes
-
Specific Outcomes:
- Established and validated the XAI Perceptual Processing Framework and the RexNet framework.
- Proposed three forms of explanations—contrastive saliency maps, counterfactual example generation, and contrastive cues—suitable for speech emotion recognition tasks.
- Experiments demonstrated that counterfactual examples and contrastive cue explanations are effective in improving user decision quality and trust.
-
Advantages Compared to Existing Solutions:
- Compared to existing explanation methods, RexNet’s multi-faceted contrastive explanations significantly enhance users' understanding of model predictions.
- Unlike traditional saliency map explanations, counterfactual examples and contrastive cues are more directly aligned with human perceptual processes.
-
Experimental and Evaluation Results:
- Experiments showed RexNet achieved an initial emotion classification accuracy of 79.5%, a significant improvement over traditional CNN models (75.7%).
- User experiments revealed that explanations combining counterfactual examples and contrastive cues effectively improved decision accuracy and trust in the model, while saliency explanations alone had limited impact.
-
Limitations and Future Directions:
- Limitations:
- Current counterfactual example generation suffers from accuracy issues.
- The interpretability of saliency information needs further refinement.
- Using instance data (rather than generated speech) limits the diversity of counterfactual examples.
- Future Research Directions:
- Employ more advanced speech synthesis techniques (e.g., Seq2Seq-VC or StarGAN-VC v2) to improve the quality of counterfactual generation.
- Explore ways to further refine the semantic content of saliency maps.
- Apply this framework to other complex domains (e.g., medical diagnosis or industrial inspection) for perceptual task explanations.
- Limitations:
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How can explainable AI be provided in speech emotion recognition that aligns with user cognition?Category: Explanation Form Design and Comprehension EffectsSimilar questionsarrow_forward
- How do contrastive explanations (e.g., saliency maps, counterfactual examples) align with human perceptual processes?Category: Explanation Form Design and Comprehension EffectsSimilar questionsarrow_forward
- How can a multimodal framework be designed to improve user understanding of and trust in deep learning model predictions?Category: Explanation Form Design and Comprehension EffectsSimilar questionsarrow_forward
lightbulb
Practical Problems
1- General users cannot understand speech emotion recognition results and hesitate to use AI.Category: Explanation Form Design and Comprehension EffectsSimilar questionsarrow_forward
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/abs/10.1145/3491102.3501826
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2022
emoji_events
Award
Best Paper
group
Authors
2 authors
sell
Subtopics
Eye Tracking & Gaze Interaction, Explainable AI (XAI)
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
0 related papers