Conversational Explanations: Discussing Explainable AI with Non-AI Experts
Authors
Explainable AI (XAI) aims to provide insights into the decisions made by AI models. To date, most XAI approaches provide only one-time, static explanations, which cannot cater to users' diverse knowledge levels and information needs. Conversational explanations have been proposed as an effective method to customize XAI explanations. However, building conversational explanation systems is hindered by the scarcity of training data. Training with synthetic data faces two main challenges: lack of data diversity and hallucination in the generated data. To alleviate these issues, we introduce a repetition penalty to promote data diversity and exploit a hallucination detector to filter out untruthful synthetic conversation turns. We conducted both automatic and human evaluations on the proposed system, fEw-shot Multi-round ConvErsational Explanation (EMCEE). For automatic evaluation, EMCEE achieves relative improvements of 81.6% in BLEU and 80.5% in ROUGE compared to the baselines. EMCEE also mitigates the degeneration of data quality caused by training on synthetic data. In human evaluations (N=60), EMCEE outperforms baseline models and the control group in improving users' comprehension, acceptance, trust, and collaboration with static explanations by large margins. Through a fine-grained analysis of model responses, we further demonstrate that training on self-generated synthetic data improves the model’s ability to generate more truthful and understandable answers, leading to better user interactions. To the best of our knowledge, this is the first conversational explanation method that can answer free-form user questions following static explanations.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can an interactive conversational XAI system be designed to meet users' knowledge levels, information needs, and task contexts?Category: Misinformation, Content Labeling, and Authenticity TrustSimilar questionsarrow_forward
- How can few-shot generation and filtering mechanisms improve diversity and quality of synthetic dialogue data?Category: Misinformation, Content Labeling, and Authenticity TrustSimilar questionsarrow_forward
- What innovative methods can reduce hallucination (misinformation) problems in generated dialogue?Category: Misinformation, Content Labeling, and Authenticity TrustSimilar questionsarrow_forward
Practical Problems
1- Non-expert users struggle to understand and trust complex AI systems' decision rationales.Category: Misinformation, Content Labeling, and Authenticity TrustSimilar questionsarrow_forward
- 75%
Manipulating and Measuring Model Interpretability
CHI '21· Explainable AI (XAI) +1
- 75%
"Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AI
CHI '21· Explainable AI (XAI) +1
- 75%
Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance
CHI '21· Explainable AI (XAI) +1
- 75%
Shared Interest: Measuring Human-AI Alignment to Identify Recurring Patterns in Model Behavior
CHI '22· Explainable AI (XAI) +1
- 75%
Debiased-CAM to mitigate image perturbations with faithful visual explanations of machine learning
CHI '22· Explainable AI (XAI) +1
- 75%
Fairness Evaluation in Text Classification: Machine Learning Practitioner Perspectives of Individual and Group Fairness
CHI '23· Explainable AI (XAI) +1
- 75%
One AI Does Not Fit All: A Cluster Analysis of the Laypeople’s Perception of AI Roles
CHI '23· Explainable AI (XAI) +1
- 75%
"Help Me Help the AI": Understanding How Explainability Can Support Human-AI Interaction
CHI '23· Explainable AI (XAI) +1
- 75%
Don’t Just Tell Me, Ask Me: AI Systems that Intelligently Frame Explanations as Questions Improve Human Logical Discernment Accuracy over Causal AI explanations
CHI '23· Explainable AI (XAI) +1
- 75%
DeepSeer: Interactive RNN Explanation and Debugging via State Abstraction
CHI '23· Explainable AI (XAI) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)