Towards Relatable Explainable AI with the Perceptual Process

Best Paper
Eye Tracking & Gaze InteractionExplainable AI (XAI)

Document Title

Towards Relatable Explainable AI with the Perceptual Process

Document Information

  • Subject Area: Explainable Artificial Intelligence (XAI), particularly contrastive explanations in perceptual tasks
  • Keywords: Explainable AI (XAI), contrastive explanations, speech emotion recognition, deep learning models, user needs, human cognition

Research Background and Issues

  • Problems and Challenges:

    • Machine learning models are often highly complex, making them difficult for end users to understand, which limits their real-world applications.
    • Current contrastive explanation methods remain overly basic, lack semantic meaning, and are difficult to align with human cognition.
    • In the audio data domain, existing explanation methods (e.g., saliency maps based on spectrograms) are overly technical and challenging for general users to comprehend.
  • Significance:

    • Explainability is a key factor in achieving responsible and trustworthy artificial intelligence.
    • Applications in the audio domain (e.g., speech emotion recognition) require explanation methods that align with human perceptual processes to enhance usability and trust.
  • Research Motivation:

    • Inspired by the theory of human perceptual processes, this study proposes a more relatable explanation framework that aligns explanations with users' conceptual, hypothetical, and associative cognitive needs.
    • It explores various contrastive explanations, including saliency, counterfactuals, and supplementary cues.

Solution

  • Proposed Framework and Model:

    • XAI Perceptual Processing Framework: Based on the theory of human perceptual processes, this framework is divided into three stages: Select, Organize, and Interpret.
    • RexNet Model (Relatable Explanation Network): A modular multi-task deep learning framework supporting three types of contrastive explanations:
      1. Contrastive Saliency: Generates contrastive saliency maps to help users understand the input features the model focuses on.
      2. Counterfactual Synthetic: Uses a generative adversarial network (StarGAN-VC) to generate style-transformed speech samples, showing differences from the target classification.
      3. Contrastive Cues: Provides emotion-related speech cues (e.g., high-frequency energy, volume) for comparison.
  • Innovations:

    • Incorporates human cognitive psychology by introducing perceptual theories into AI explanation design.
    • Proposes a model architecture that organically integrates saliency, counterfactual examples, and cue information.
    • First application of relatable explanation methods in audio prediction tasks.
  • Implementation Steps:

    • Train a speech emotion classification model using the RAVDESS dataset and extend it into RexNet, achieving emotion prediction through a spectrogram-based CNN.
    • Integrate explanation modules and generate explanations using techniques such as Grad-CAM, generative adversarial networks (GAN), and layer-wise relevance propagation (LRP).
    • Evaluate the usability and effectiveness of the explanations through modeling studies and user experiments.

Research Outcomes

  • Specific Outcomes:

    • Established and validated the XAI Perceptual Processing Framework and the RexNet framework.
    • Proposed three forms of explanations—contrastive saliency maps, counterfactual example generation, and contrastive cues—suitable for speech emotion recognition tasks.
    • Experiments demonstrated that counterfactual examples and contrastive cue explanations are effective in improving user decision quality and trust.
  • Advantages Compared to Existing Solutions:

    • Compared to existing explanation methods, RexNet’s multi-faceted contrastive explanations significantly enhance users' understanding of model predictions.
    • Unlike traditional saliency map explanations, counterfactual examples and contrastive cues are more directly aligned with human perceptual processes.
  • Experimental and Evaluation Results:

    • Experiments showed RexNet achieved an initial emotion classification accuracy of 79.5%, a significant improvement over traditional CNN models (75.7%).
    • User experiments revealed that explanations combining counterfactual examples and contrastive cues effectively improved decision accuracy and trust in the model, while saliency explanations alone had limited impact.
  • Limitations and Future Directions:

    • Limitations:
      • Current counterfactual example generation suffers from accuracy issues.
      • The interpretability of saliency information needs further refinement.
      • Using instance data (rather than generated speech) limits the diversity of counterfactual examples.
    • Future Research Directions:
      • Employ more advanced speech synthesis techniques (e.g., Seq2Seq-VC or StarGAN-VC v2) to improve the quality of counterfactual generation.
      • Explore ways to further refine the semantic content of saliency maps.
      • Apply this framework to other complex domains (e.g., medical diagnosis or industrial inspection) for perceptual task explanations.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/68763/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/abs/10.1145/3491102.3501826
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2022
emoji_events
Award
Best Paper
group
Authors
2 authors
sell
Subtopics
Eye Tracking & Gaze Interaction, Explainable AI (XAI)
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
0 related papers