Unraveling the Dilemma of AI Errors: Exploring the Effectiveness of Human and Machine Explanations for Large Language Models
Authors
Title of the Paper
Deciphering the Dilemma of AI Errors: Exploring the Effectiveness of Human and Machine Interpretations in Large Language Models
Paper Information
- Field of Study: Explainable Artificial Intelligence (XAI), particularly its application in large language models
- Keywords: Explainable Artificial Intelligence (XAI), Large Language Models (LLMs), saliency maps, textual explanations, post-hoc explanations, Stanford Question Answering Dataset (SQuAD 1.1v), user studies, explanation bias, evaluation of explanations, human explanations, machine explanations
Research Background and Issues
-
Issues or Challenges:
- With the rise of deep learning and large language models, AI has become an integral part of human life, making the need for explainability increasingly urgent.
- Current explanation methods (e.g., saliency maps) provide visualizations beyond model predictions, but their long-term effectiveness and applicability, especially in cases of incorrect model outputs, remain questionable.
- There is skepticism about the use of existing visualization methods (e.g., saliency maps), and XAI approaches lack comprehensive controlled conditions and human-participatory studies.
-
Significance: AI decisions can impact critical societal domains such as medical diagnoses or loan approvals. Transparency and explainability are not only technical requirements but also legal mandates (e.g., explanation provisions in GDPR).
-
Research Motivation and Related Work:
- Deep learning models are often criticized as "black boxes," making it difficult to explain individual predictions or generated textual content.
- The differences between human and machine explanations and their impact on performance, satisfaction, and trust remain unclear.
- A deeper comparison between human and machine explanations is needed, especially in cases involving incorrect predictions.
Proposed Solution
-
Proposed Approach:
- Collect 156 human-generated saliency maps and textual explanations, and compare them with state-of-the-art machine explanations (e.g., Conservative Layer-wise Relevance Propagation (Conservative LRP), Integrated Gradients (IG), and ChatGPT-generated explanations).
- Design a three-part experiment, including explanation collection, analysis, and human participant studies to evaluate the effectiveness of explanations.
-
Innovations:
- Conduct the first evaluation of saliency maps in text generation tasks, considering both correct and incorrect model predictions.
- Provide a direct comparison between human-generated and machine-generated explanations, using quantitative and qualitative evaluations to understand their impact on user trust, satisfaction, and task performance.
- Introduce the concept of "explanation confirmation bias" and explore the dilemma of explanations related to incorrect predictions.
-
Implementation Steps and Key Techniques:
- Explanation Collection and Generation:
- Use question-answering task samples from the SQuAD 1.1v dataset, combining human-generated saliency maps and textual explanations.
- Generate machine saliency maps using existing techniques (Conservative LRP and IG) and textual explanations using ChatGPT.
- Explanation Analysis:
- Conduct thematic analysis to classify human textual explanations into groups such as "extraction," "interpretation," "misunderstanding," and "poor quality."
- Perform quantitative evaluations of the overlap between human and machine-generated saliency maps.
- Human Studies:
- Compare and evaluate subjective factors such as satisfaction, trust, quality, and helpfulness of human and machine explanations, while tracking objective task performance (accuracy and task time).
- Explanation Collection and Generation:
Research Findings
-
Specific Results:
- Human-generated saliency maps and controlled condition explanations (e.g., displaying only the answer location) were deemed more helpful than machine-generated saliency maps, with shorter task completion times.
- In textual explanations, participants showed significantly higher trust in textual extraction compared to ChatGPT-generated explanations.
- The correctness of AI predictions had a more significant impact on task performance and subjective perceptions than the type of explanation (affecting cognitive load, time investment, satisfaction, etc.).
- Identified the dilemma of explanations: "Effective explanations supporting incorrect predictions may reduce task performance."
-
Advantages Over Existing Solutions:
- Compared multiple human and machine-generated explanations, focusing on real-world task scenarios and human participation.
- Captured how participants exhibited trust mechanisms and cognitive biases when faced with explanations, particularly for incorrect answers.
-
Experimental Results:
- Analyzed subjective ratings (e.g., satisfaction, trust) and objective metrics (e.g., task time, accuracy).
- In cases of incorrect answers, satisfaction and trust were negatively correlated with task accuracy.
-
Limitations and Future Directions:
- The tasks in the study were relatively short, and the exploration capability of saliency maps in long texts was not fully tested.
- Human participants lacked specialized training on what constitutes a good explanation and could not directly compare different explanations.
- Enhancing AI interactive design and providing related case studies may effectively reduce the impact of incorrect explanations in the future.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- Which is more effective for understanding LLM outputs: human-generated explanations or machine-generated explanations?Category: Explanation Form Design and Comprehension EffectsSimilar questionsarrow_forward
- Do explanations in cases of incorrect predictions reduce users' task performance?Category: Explanation Form Design and Comprehension EffectsSimilar questionsarrow_forward
- How does 'explanation confirmation bias' affect user trust and cognitive load?Category: Explanation Form Design and Comprehension EffectsSimilar questionsarrow_forward
Practical Problems
1- Users struggle to understand AI decisions, especially incorrect predictions, affecting task performance.Category: Explanation Form Design and Comprehension EffectsSimilar questionsarrow_forward
- 83%
The Disagreement Deconvolution: Bringing Machine Learning Performance Metrics In Line With Reality
CHI '21· Explainable AI (XAI) +1
- 71%
Evaluating the Interpretability of Generative Models by Interactive Reconstruction
CHI '21· Explainable AI (XAI) +2
- 71%
LLM-box vs. Thinking-box: Designing for Deliberate User Engagement with Distorted Information in Conversational Search
CHI '26· Human-LLM Collaboration +2
- 71%
Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving Tasks
CHI '26· Human-LLM Collaboration +2
- 71%
The Impact of Response Latency and Task Type on Human-LLM Interaction and Perception
CHI '26· Human-LLM Collaboration +2
- 71%
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
CHI '26· Human-LLM Collaboration +2
- 71%
Exploring the Innovation Opportunities for Pre-trained Models
DIS '25· Generative AI (Text, Image, Music, Video) +2
- 67%
Manipulating and Measuring Model Interpretability
CHI '21· Explainable AI (XAI) +1
- 67%
Shared Interest: Measuring Human-AI Alignment to Identify Recurring Patterns in Model Behavior
CHI '22· Explainable AI (XAI) +1
- 67%
Debiased-CAM to mitigate image perturbations with faithful visual explanations of machine learning
CHI '22· Explainable AI (XAI) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)